Alibaba launches Qwen-Audio-3.0-TTS for efficient, robust, and consistent text-to-speech

Alibaba launches Qwen-Audio-3.0-TTS for efficient, robust, and consistent text-to-speech

Alibaba Cloud has released Qwen-Audio-3.0-TTS, a production-oriented speech synthesis system, bringing major technical advances for users requiring high-quality, customizable text-to-speech. The latest version enhances content consistency, speaker similarity, prosodic naturalness, audio quality, controllability, multilingual coverage, overall efficiency, and robustness.

Building on these improvements, the system integrates a 12.5 Hz low-frame-rate speech tokenizer that lowers inference latency, alongside a five-stage progressive training paradigm that tightly coordinates language and frequency modeling optimization. Qwen-Audio-3.0-TTS also introduces free-style natural-language instruction following and fine-grained inline tagging, enabling production-level control for a range of deployment scenarios.

For global users, Qwen-Audio-3.0-TTS supports 16 languages, 20 Chinese dialect regions, and offers one-pass long-form synthesis for outputs up to 3 minutes. It maintains high performance even when generating speech from noisy, reverberant, or unclear reference audio sources.

Following these advances, Alibaba’s model achieves state-of-the-art results across multiple evaluations, including SEED-TTS-Eval, CV3-Eval, instruction-following, long-form, and acoustic-robustness tests. It also claims the top position on the independent Artificial Analysis Text-to-Speech Leaderboard.

by Paul

Add as a preferred source on Google
  • ...

Qwen Audio is a multimodal model designed for processing both audio and text. It supports over 30 language and sound tasks, featuring capabilities such as AI Voice Cloning and Speech Recognition. Qwen Audio excels in multi-turn dialogue and universal audio understanding, demonstrating strong performance in benchmark tests.

No comments so far, maybe you want to be first?
Gu