Qwen Audio icon
Qwen Audio icon

Qwen Audio

Processes both audio and text inputs with a multi-task framework supporting over 30 language and sound tasks, enabling multi-turn dialogue, sound reasoning, and tool use, while excelling in benchmarks without task-specific fine-tuning or retraining.

Qwen Audio screenshot 1

Cost / License

Platforms

  • Mac
  • Windows
  • Linux
  • FFmpeg
  • Python
  • PyTorch
0likes
0comments
0articles

Features

Qwen Audio News & Activities

Highlights All activities

Recent News

Recent activities

Qwen Audio information

  • Developed by

    CN flagAlibaba Cloud
  • Licensing

    Open Source (Apache-2.0) and Free product.
  • Written in

  • Alternatives

    51 alternatives listed
  • Supported Languages

    • English

AlternativeTo Category

AI Tools & Services

GitHub repository

  •  1,914 Stars
  •  145 Forks
  •  63 Open Issues
  •   Updated  
View on GitHub
Qwen Audio was added to AlternativeTo by Paul on and this page was last updated .
No comments or reviews, maybe you want to be first?

What is Qwen Audio?

Qwen-Audio (Qwen Large Audio Language Model) is the multimodal version of the large model series, Qwen (abbr. Tongyi Qianwen), proposed by Alibaba Cloud. Qwen-Audio accepts diverse audio (human speech, natural sound, music and song) and text as inputs, outputs text. The contribution of Qwen-Audio include:

  • Fundamental audio models: Qwen-Audio is a fundamental multi-task audio-language model that supports various tasks, languages, and audio types, serving as a universal audio understanding model. Building upon Qwen-Audio, we develop Qwen-Audio-Chat through instruction fine-tuning, enabling multi-turn dialogues and supporting diverse audio-oriented scenarios.
  • Multi-task learning framework for all types of audios: To scale up audio-language pre-training, we address the challenge of variation in textual labels associated with different datasets by proposing a multi-task training framework, enabling knowledge sharing and avoiding one-to-many interference. Our model incorporates more than 30 tasks and extensive experiments show the model achieves strong performance.
  • Strong Performance: Experimental results show that Qwen-Audio achieves impressive performance across diverse benchmark tasks without requiring any task-specific fine-tuning, surpassing its counterparts. Specifically, Qwen-Audio achieves state-of-the-art results on the test set of Aishell1, cochlscene, ClothoAQA, and VocalSound.
  • Flexible multi-run chat from audio and text input: Qwen-Audio supports multiple-audio analysis, sound understanding and reasoning, music appreciation, and tool usage.

Official Links