
Alibaba launches Qwen-Audio-3.0-TTS for efficient, robust, and consistent text-to-speech
Alibaba Cloud has released Qwen-Audio-3.0-TTS, a production-oriented speech synthesis system, bringing major technical advances for users requiring high-quality, customizable text-to-speech. The latest version enhances content consistency, speaker similarity, prosodic naturalness, audio quality, controllability, multilingual coverage, overall efficiency, and robustness.
Building on these improvements, the system integrates a 12.5 Hz low-frame-rate speech tokenizer that lowers inference latency, alongside a five-stage progressive training paradigm that tightly coordinates language and frequency modeling optimization. Qwen-Audio-3.0-TTS also introduces free-style natural-language instruction following and fine-grained inline tagging, enabling production-level control for a range of deployment scenarios.
For global users, Qwen-Audio-3.0-TTS supports 16 languages, 20 Chinese dialect regions, and offers one-pass long-form synthesis for outputs up to 3 minutes. It maintains high performance even when generating speech from noisy, reverberant, or unclear reference audio sources.
Following these advances, Alibaba’s model achieves state-of-the-art results across multiple evaluations, including SEED-TTS-Eval, CV3-Eval, instruction-following, long-form, and acoustic-robustness tests. It also claims the top position on the independent Artificial Analysis Text-to-Speech Leaderboard.


