Alibaba's Qwen-Audio-3.0-TTS Tops Artificial Analysis Ranking
The Plus model's top ranking on an independent TTS leaderboard, combined with Japanese support and improved voice cloning from imperfect audio, positions Alibaba's open release as a direct competitor in the text-to-speech space.
Reporting from 1 source: GameBusiness.jp.
Alibaba's Tongyi Lab released Qwen-Audio-3.0-TTS, a text-to-speech model with two variants. The Plus model ranks first on Artificial Analysis's TTS leaderboard. It supports 16 languages including Japanese, offers natural-language control over emotion and pacing, and can clone voices from noisy reference audio.
Tongyi Lab, the Alibaba research unit behind the Qwen series, has released Qwen-Audio-3.0-TTS, a text-to-speech model that converts input text into speech and can reproduce a specific person's voice from a reference audio clip.
The release includes two models. Flash is optimized for real-time interaction, outputting the first audio data in roughly 300 milliseconds. Plus prioritizes naturalness and timbre fidelity over speed, and it has taken first place on Artificial Analysis's TTS leaderboard, an independent third-party ranking.
The model supports 16 languages, including Japanese. Users can control emotion, character settings, scenario, and speaking pace through natural-language instructions, and can insert tags directly into the target text to specify non-verbal details like breaths, laughter, and tone shifts. The voice cloning function also improves, automatically suppressing noise from imperfect reference audio while preserving the original voice quality.
Synthesized by Yomimono from the 1 cited source below, including Japanese-language reporting where cited, then editorially reviewed before publishing.