{"id":64162,"date":"2026-07-21T01:20:51","date_gmt":"2026-07-21T05:20:51","guid":{"rendered":"https:\/\/overcentral.com\/en\/?p=64162"},"modified":"2026-07-21T01:20:51","modified_gmt":"2026-07-21T05:20:51","slug":"qwen-audio-3-0-tts","status":"publish","type":"post","link":"https:\/\/overcentral.com\/en\/qwen-audio-3-0-tts\/","title":{"rendered":"Qwen-Audio-3.0-TTS Launches in Flash and Plus Tiers Across 16 Languages"},"content":{"rendered":"<p>Alibaba&#8217;s Tongyi Lab has launched <strong>Qwen-Audio-3.0-TTS<\/strong>, a production-oriented text-to-speech system that ships in two distinct tiers \u2014 Flash and Plus \u2014 covering 16 languages. The model is delivered exclusively as a hosted service through <a href=\"https:\/\/www.aliyun.com\/product\/modelstudio\" target=\"_blank\" rel=\"noopener noreferrer\" data-iacss-external=\"1\">Alibaba Cloud Model Studio<\/a> rather than as downloadable weights, marking a clear bet on API-first deployment for voice synthesis workloads.<\/p>\n<h2>Two tiers, one model family<\/h2>\n<p>Flash is optimized for real-time interaction, with first-packet latency around 300 milliseconds, making it suitable for voice agents and live assistants where response speed is critical. Plus prioritizes generation quality, targeting audiobooks, narration, and dubbing where naturalness and timbre fidelity matter more than raw speed. Both variants share the same underlying architecture but are tuned for different points on the latency-quality curve.<\/p>\n<p>The API is accessed over a bidirectional WebSocket streaming protocol and supports PCM, WAV, MP3, and Opus output at sample rates up to 48 kHz. Alibaba provides the DashScope SDK alongside raw WebSocket examples in Python, Java, Go, C#, PHP, and Node.js, with service regions in Singapore and Beijing.<\/p>\n<h2>Technical foundation<\/h2>\n<p>Two design decisions anchor the system. A 12.5 Hz low-frame-rate speech tokenizer reduces autoregressive decoding cost by generating fewer tokens per second of audio, directly cutting inference latency. A five-stage progressive training pipeline coordinates the language model and flow-matching components through independent pretraining, joint high-quality data annealing, and separate reinforcement learning stages for the LM and the flow-matching module. The research team reports that this pipeline improves content consistency, prosodic naturalness, voice fidelity, perceptual quality, and robustness against noisy reference audio.<\/p>\n<p>The model also handles one-pass long-form synthesis up to three minutes, hard text-normalization cases, and vocoder super-resolution for 48 kHz output.<\/p>\n<h2>Multilingual coverage across 16 languages<\/h2>\n<p>Qwen-Audio-3.0-TTS supports Arabic, Chinese, English, French, German, Indonesian, Italian, Japanese, Korean, Malay, Portuguese, Russian, Spanish, Tagalog, Thai, and Vietnamese. Seven of these languages are new additions compared to the prior CosyVoice-3.0 line. The model also covers 20 Chinese dialect regions.<\/p>\n<p>On multilingual intelligibility, the family achieves the best word or character error rate in 10 of the 16 languages. Flash delivers the lowest average WER at 3.87, while Plus follows closely at 3.96. On speaker similarity, Plus ranks first across all 16 languages with an average of 82.75, with Flash at 80.44. A curated preset voice library spanning all 16 supported languages is also included, so teams can ship a voice without cloning one first.<\/p>\n<h2>Inline tag control for fine-grained expression<\/h2>\n<p>The system embeds 86 inline tags directly inside the target text for localized control at the phrase and word level. These are split into two categories. Control tags such as <code>[excited]<\/code>codecodecode, <code>[sad]<\/code>codecodecode, <code>[whispers]<\/code>codecodecode, and <code>[asmr]<\/code>codecodecode set an emotion or style that persists until the next tag. Rich-language tags such as <code>[laughing]<\/code>codecodecode, <code>[gasp]<\/code>codecodecode, and <code>[clears throat]<\/code>codecodecode insert a single vocal effect without changing the surrounding tone. A worked example from the documentation: <code>[excited]What a beautiful day today![laughing]Let's go out and have fun together!<\/code>codecodecode These emotion and rich-language tags are supported only in unidirectional streaming mode.<\/p>\n<h2>Leaderboard standing and trade-offs<\/h2>\n<p>Qwen-Audio-3.0-TTS-Plus took the top quality spot on the Artificial Analysis Speech Arena for Provider Voices, posting an Elo near 1,236, narrowly ahead of Simba 3.2 at 1,234 and clear of <a href=\"https:\/\/overcentral.com\/en\/google-home-speaker-gemini-review\/\" title=\"Google Home Speaker launches with unfinished Gemini for Home\" data-iacss-internal=\"1\">Gemini<\/a> 3.1 Flash TTS at 1,214 and Sonic 3.5 at 1,207. The lead over Simba 3.2 sits inside overlapping confidence intervals, effectively a statistical tie at the very top.<\/p>\n<p>Two trade-offs are worth stating plainly. Throughput is modest: Plus generates about 16 characters per second, below Simba 3.2 at 30.2 and Sonic 3.5 at 120. Price is competitive: the listed rate is $27.59 per 1 million characters, roughly a third of what ElevenLabs and MiniMax charge for the tiers it outranks. <a href=\"https:\/\/overcentral.com\/en\/rezero-anime-corner-rankings\/\" title=\"Re:Zero Dominates Anime Corner Spring 2026 Rankings\" data-iacss-internal=\"1\">Rankings<\/a> and pricing shift frequently, so confirmation is advisable before planning around them.<\/p>\n<h2>What this means for developers<\/h2>\n<p>Qwen-Audio-3.0-TTS gives development teams a production-grade hosted TTS option that leads the independent quality leaderboard at a significantly lower price point than Western incumbents. The Flash tier is a strong candidate for real-time voice agents, while the Plus tier suits high-fidelity content generation. The key constraint is that the model is API-only \u2014 teams that require self-hosted or offline inference should look to the open-weight Qwen3-TTS line, which is Apache-2.0 licensed and architecturally distinct. Developers can begin experimenting immediately through Alibaba Cloud Model Studio using the provided SDKs and WebSocket examples across six programming languages.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Alibaba&#8217;s Tongyi Lab has launched Qwen-Audio-3.0-TTS, a production-oriented text-to-speech system that ships in two distinct tiers \u2014 Flash and Plus \u2014 covering 16 languages. The model is delivered exclusively as a hosted service through Alibaba Cloud Model Studio rather than as downloadable weights, marking a clear bet on API-first deployment for voice synthesis workloads. Two [&hellip;]<\/p>\n","protected":false},"author":7,"featured_media":83798,"comment_status":"closed","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"fifu_image_url":"https:\/\/cards.overcentral.com\/cards\/en\/64162.png","fifu_image_alt":"Qwen-Audio-3.0-TTS Launches in Flash and Plus Tiers Across 16 Languages","footnotes":""},"categories":[349],"tags":[],"class_list":["post-64162","post","type-post","status-publish","format-standard","has-post-thumbnail","category-articles"],"fifu_image_url":"https:\/\/cards.overcentral.com\/cards\/en\/64162.png","fifu_image_alt":"Qwen-Audio-3.0-TTS Launches in Flash and Plus Tiers Across 16 Languages","_links":{"self":[{"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/posts\/64162","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/users\/7"}],"replies":[{"embeddable":true,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/comments?post=64162"}],"version-history":[{"count":0,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/posts\/64162\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/media\/83798"}],"wp:attachment":[{"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/media?parent=64162"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/categories?post=64162"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/tags?post=64162"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}