• bitcoinBitcoin(BTC)$66,117.002.75%
  • ethereumEthereum(ETH)$1,930.053.02%
  • tetherTether(USDT)$1.000.03%
  • binancecoinBNB(BNB)$576.851.77%
  • usd-coinUSDC(USDC)$1.000.01%
  • rippleXRP(XRP)$1.133.32%
  • solanaSolana(SOL)$78.112.15%
  • tronTRON(TRX)$0.3270490.37%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.00-1.68%
  • HyperliquidHyperliquid(HYPE)$62.532.81%
  • dogecoinDogecoin(DOGE)$0.0734621.86%
  • USDSUSDS(USDS)$1.000.01%
  • RainRain(RAIN)$0.013913-2.10%
  • zcashZcash(ZEC)$536.730.76%
  • leo-tokenLEO Token(LEO)$9.710.40%
  • whitebitWhiteBIT Coin(WBT)$57.692.68%
  • stellarStellar(XLM)$0.1916322.74%
  • cardanoCardano(ADA)$0.1745477.01%
  • chainlinkChainlink(LINK)$8.693.49%
  • moneroMonero(XMR)$346.263.39%
  • CantonCanton(CC)$0.1259101.32%
  • daiDai(DAI)$1.000.00%
  • bitcoin-cashBitcoin Cash(BCH)$224.275.44%
  • USD1USD1(USD1)$1.000.03%
  • Ethena USDeEthena USDe(USDE)$1.000.01%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.451.23%
  • litecoinLitecoin(LTC)$47.300.52%
  • Global DollarGlobal Dollar(USDG)$1.00-0.27%
  • suiSui(SUI)$0.773.17%
  • hedera-hashgraphHedera(HBAR)$0.0683303.48%
  • Circle USYCCircle USYC(USYC)$1.13-0.01%
  • avalanche-2Avalanche(AVAX)$6.621.28%
  • paypal-usdPayPal USD(PYUSD)$1.000.01%
  • crypto-com-chainCronos(CRO)$0.0579720.71%
  • nearNEAR Protocol(NEAR)$1.992.47%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • shiba-inuShiba Inu(SHIB)$0.0000041.47%
  • tether-goldTether Gold(XAUT)$4,053.820.89%
  • uniswapUniswap(UNI)$3.696.79%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.140.28%
  • OndoOndo(ONDO)$0.40430815.90%
  • BittensorBittensor(TAO)$199.602.45%
  • pax-goldPAX Gold(PAXG)$4,050.780.87%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.056596-0.48%
  • okbOKB(OKB)$81.841.92%
  • AsterAster(ASTER)$0.631.36%
  • HTX DAOHTX DAO(HTX)$0.0000020.54%
  • Ripple USDRipple USD(RLUSD)$1.000.01%
  • MemeCoreMemeCore(M)$1.17-5.08%
  • usddUSDD(USDD)$1.00-0.01%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Alibaba’s Tongyi Lab Releases Qwen-Audio-3.0-TTS, a Hosted Text-to-Speech Model in Flash and Plus Tiers Across 16 Languages

July 20, 2026
in AI & Technology
Reading Time: 23 mins read
A A
Alibaba’s Tongyi Lab Releases Qwen-Audio-3.0-TTS, a Hosted Text-to-Speech Model in Flash and Plus Tiers Across 16 Languages
ShareShareShareShareShare

Alibaba’s Tongyi Lab has released Qwen-Audio-3.0-TTS, a production-oriented text-to-speech (TTS) system. The model ships in two variants from the same lineage. Flash targets real-time interaction. Plus targets high-quality generation. Both are delivered as hosted models through Alibaba Cloud Model Studio, not as downloadable weights.

The release focuses on four things developers hit in production: broader language coverage, natural-language style control, fine-grained tag control, and robustness when the reference audio is not clean. Qwen-Audio-3.0-TTS-Plus also ranks first on the independent Artificial Analysis Text-to-Speech leaderboard.

YOU MAY ALSO LIKE

Amazon’s Adaptive Display For Fire TVs Is Officially Rolling Out Today

The First UL 3700-Compliant Plug-In Solar Microinverter Is Now Available In The US

Two variants, one lineage

The two tiers map to different jobs. Flash is tuned for real-time interaction, with first-packet latency at the 300 ms level. Plus is tuned for high-quality generation, where naturalness and timbre fidelity matter more than speed.

The model IDs are qwen-audio-3.0-tts-flash and qwen-audio-3.0-tts-plus. Both are called over a bidirectional WebSocket streaming protocol. The API supports PCM, WAV, MP3, and Opus, with sample-rate output up to 48 kHz. It exposes streaming input and output, voice cloning, Voice Design, and instruction control. Alibaba provides the DashScope SDK plus raw WebSocket examples in Python, Java, Go, C#, PHP, and Node.js, across its Singapore and Beijing regions.

How the model is built

Two design choices anchor the system.

  1. A 12.5 Hz low-frame-rate speech tokenizer reduces autoregressive decoding cost while retaining content and speaker information. A lower frame rate means fewer tokens per second of audio, which cuts inference latency.
  2. A five-stage progressive training paradigm coordinates the language model (LM) and flow-matching (FM) components. The stages are independent LM and FM pretraining, joint training with high-quality data annealing, LM reinforcement learning, FM robustness training, and FM reinforcement learning. The research team reports this pipeline improves content consistency, prosodic naturalness, voice fidelity, perceptual quality, and robustness.

The model also handles one-pass long-form synthesis up to 3 minutes, hard text-normalization cases, and vocoder super-resolution for 48 kHz output.


Multilingual coverage across 16 languages

Qwen-Audio-3.0-TTS supports 16 languages: Arabic, Chinese, English, French, German, Indonesian, Italian, Japanese, Korean, Malay, Portuguese, Russian, Spanish, Tagalog, Thai, and Vietnamese. Seven of these are newly added versus the prior line. It also covers 20 Chinese dialect regions.

On multilingual intelligibility, the model family posts the best word/character error rate (WER/CER) in 10 of the 16 languages. Flash delivers the lowest average WER/CER at 3.87; Plus is close at 3.96. Lower is better on this metric.

On speaker similarity, Plus ranks first across all 16 languages with an average of 82.75, and Flash follows at 80.44. The release also adds a curated preset voice library spanning the 16 supported languages, so teams can ship a voice without cloning one first.

For precise control, the research team embed inline tags directly in the target text. The release adds 86 fine-grained inline tags for localized control at the phrase and word level. These cover expressive transitions and non-verbal events such as laughter, breathing, coughing, and sighing.

The Model Studio documentation splits these into two groups. Control tags such as [excited], [sad], [whispers], and [asmr] set an emotion or style until the next tag. Rich-language tags such as [laughing], [gasp], and [clears throat] insert a single vocal effect without changing surrounding tone. A worked example: [excited]What a beautiful day today![laughing]Let's go out and have fun together! One limitation is worth noting here: these emotion and rich-language tags are supported only in unidirectional streaming mode.

Where it stands on the leaderboard

Qwen-Audio-3.0-TTS-Plus took the top quality spot on the Artificial Analysis Speech Arena for Provider Voices. It posts an Elo near 1,236, narrowly ahead of Simba 3.2 at 1,234, and clear of Gemini 3.1 Flash TTS (1,214) and Sonic 3.5 (1,207). The lead over Simba 3.2 sits inside overlapping confidence intervals, so it is a statistical tie at the very top.

Two trade-offs are worth stating plainly. Throughput is modest: Plus generates about 16 characters per second, below Simba 3.2 (30.2), Gemini 3.1 Flash TTS (27), and Sonic 3.5 (120). Price is competitive: the listed rate is $27.59 per 1M characters, roughly a third of what ElevenLabs and MiniMax charge for the tiers it outranks. Rank and price move often, so confirm both before planning around them.

How developers and the community are reacting

The early reception is cautiously enthusiastic. The most-shared story is that a non-Western TTS topped the arena at a fraction of incumbent pricing. The most common reservations are that the model is hosted-only, that its throughput trails rivals, and that its name overlaps with the open Qwen3-TTS line. The dashboard below aggregates that early signal across X, Reddit, and Hacker News.


Key Takeaways

  • Qwen-Audio-3.0-TTS is a hosted TTS model in two tiers: Flash (~300 ms first-packet, real-time) and Plus (quality-first).
  • Plus ranks #1 on the Artificial Analysis arena at ~1,236 Elo, priced at ~$27.59 per 1M characters, but only ~16 chars/sec throughput.
  • Coverage spans 16 languages and 20 Chinese dialects, with best WER/CER in 10 of 16 languages and top speaker similarity on Plus.
  • Control comes two ways: free-style natural-language instructions and 86 fine-grained inline tags for non-verbal detail.
  • It is distinct from the open-weight Qwen3-TTS (Apache-2.0) line; the 3.0 model is API-only via Alibaba Cloud Model Studio.

Sources: Qwen-Audio-3.0-TTS release blog · Tongyi Lab announcement · Alibaba Cloud Model Studio real-time TTS docs · Artificial Analysis leaderboard · Tongyi Lab on X


Asif Razzaq is the CEO of Marktechpost Media Inc.. As a visionary entrepreneur and engineer, Asif is committed to harnessing the potential of Artificial Intelligence for social good. His most recent endeavor is the launch of an Artificial Intelligence Media Platform, Marktechpost, which stands out for its in-depth coverage of machine learning and deep learning news that is both technically sound and easily understandable by a wide audience. The platform boasts of over 2 million monthly views, illustrating its popularity among audiences.

Credit: Source link

ShareTweetSendSharePin

Related Posts

Amazon’s Adaptive Display For Fire TVs Is Officially Rolling Out Today
AI & Technology

Amazon’s Adaptive Display For Fire TVs Is Officially Rolling Out Today

July 20, 2026
The First UL 3700-Compliant Plug-In Solar Microinverter Is Now Available In The US
AI & Technology

The First UL 3700-Compliant Plug-In Solar Microinverter Is Now Available In The US

July 20, 2026
Writer’s AI harness cuts token spend nearly 40% — without sacrificing accuracy
AI & Technology

Writer’s AI harness cuts token spend nearly 40% — without sacrificing accuracy

July 20, 2026
A single AI agent conversation can look perfect and still be broken, leaders from LangChain, Conviva and CoreWeave said at VB Transform 2026
AI & Technology

A single AI agent conversation can look perfect and still be broken, leaders from LangChain, Conviva and CoreWeave said at VB Transform 2026

July 20, 2026
Next Post
Writer’s AI harness cuts token spend nearly 40% — without sacrificing accuracy

Writer's AI harness cuts token spend nearly 40% — without sacrificing accuracy

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
How to Start Investing at Schwab

How to Start Investing at Schwab

July 15, 2026
Atea ASA (ATEAY) Q2 2026 Earnings Call Transcript

Atea ASA (ATEAY) Q2 2026 Earnings Call Transcript

July 15, 2026
Amazon’s Adaptive Display For Fire TVs Is Officially Rolling Out Today

Amazon’s Adaptive Display For Fire TVs Is Officially Rolling Out Today

July 20, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!