

Qwen Audio
Like
Processes both audio and text inputs with a multi-task framework supporting over 30 language and sound tasks, enabling multi-turn dialogue, sound reasoning, and tool use, while excelling in benchmarks without task-specific fine-tuning or retraining.
Cost / License
- Free
- Open Source (Apache-2.0)
Application types
Platforms
- Mac
- Windows
- Linux
- FFmpeg
- Python
- PyTorch
Features
- Text to Speech
- AI Voice Cloning
- Speech Recognition
Qwen Audio News & Activities
Highlights All activities
Recent News
- POX published news article about Qwen Audio
Alibaba launches Qwen-Audio-3.0-TTS for efficient, robust, and consistent text-to-speechAlibaba Cloud has released Qwen-Audio-3.0-TTS, a production-oriented speech synthesis system, bring...
Recent activities
- POX updated Qwen Audio
- POX added Qwen Audio as alternative to Voxtral, VoiceCraft, X to Voice and SherpaTTS
- POX added Qwen Audio
Qwen Audio information
No comments or reviews, maybe you want to be first?
What is Qwen Audio?
Qwen-Audio (Qwen Large Audio Language Model) is the multimodal version of the large model series, Qwen (abbr. Tongyi Qianwen), proposed by Alibaba Cloud. Qwen-Audio accepts diverse audio (human speech, natural sound, music and song) and text as inputs, outputs text. The contribution of Qwen-Audio include:
- Fundamental audio models: Qwen-Audio is a fundamental multi-task audio-language model that supports various tasks, languages, and audio types, serving as a universal audio understanding model. Building upon Qwen-Audio, we develop Qwen-Audio-Chat through instruction fine-tuning, enabling multi-turn dialogues and supporting diverse audio-oriented scenarios.
- Multi-task learning framework for all types of audios: To scale up audio-language pre-training, we address the challenge of variation in textual labels associated with different datasets by proposing a multi-task training framework, enabling knowledge sharing and avoiding one-to-many interference. Our model incorporates more than 30 tasks and extensive experiments show the model achieves strong performance.
- Strong Performance: Experimental results show that Qwen-Audio achieves impressive performance across diverse benchmark tasks without requiring any task-specific fine-tuning, surpassing its counterparts. Specifically, Qwen-Audio achieves state-of-the-art results on the test set of Aishell1, cochlscene, ClothoAQA, and VocalSound.
- Flexible multi-run chat from audio and text input: Qwen-Audio supports multiple-audio analysis, sound understanding and reasoning, music appreciation, and tool usage.




