• bitcoinBitcoin(BTC)$84,608.000.62%
  • ethereumEthereum(ETH)$2,688.120.33%
  • tetherTether(USDT)$1.000.00%
  • binancecoinBNB(BNB)$777.820.93%
  • rippleXRP(XRP)$1.520.44%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$122.881.76%
  • tronTRON(TRX)$0.333637-0.45%
  • zcashZcash(ZEC)$1,605.530.93%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.062.87%
  • HyperliquidHyperliquid(HYPE)$91.810.14%
  • dogecoinDogecoin(DOGE)$0.0971371.03%
  • chainlinkChainlink(LINK)$14.050.37%
  • moneroMonero(XMR)$547.57-0.85%
  • whitebitWhiteBIT Coin(WBT)$84.380.62%
  • USDSUSDS(USDS)$1.00-0.01%
  • cardanoCardano(ADA)$0.2556691.72%
  • RainRain(RAIN)$0.012566-2.22%
  • leo-tokenLEO Token(LEO)$9.050.71%
  • stellarStellar(XLM)$0.2166280.47%
  • nearNEAR Protocol(NEAR)$5.4814.18%
  • bitcoin-cashBitcoin Cash(BCH)$334.35-0.16%
  • uniswapUniswap(UNI)$9.722.12%
  • litecoinLitecoin(LTC)$71.26-0.33%
  • CantonCanton(CC)$0.1379533.16%
  • suiSui(SUI)$1.2711.11%
  • Ethena USDeEthena USDe(USDE)$1.000.02%
  • avalanche-2Avalanche(AVAX)$10.962.52%
  • daiDai(DAI)$1.000.00%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.633.89%
  • USD1USD1(USD1)$1.00-0.01%
  • hedera-hashgraphHedera(HBAR)$0.0948572.39%
  • BittensorBittensor(TAO)$326.233.42%
  • shiba-inuShiba Inu(SHIB)$0.0000060.50%
  • crypto-com-chainCronos(CRO)$0.0672582.19%
  • BitwayBitway(BTW)$1.2116.29%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • quant-networkQuant(QNT)$198.9762.59%
  • EthenaEthena(ENA)$0.2801764.71%
  • paypal-usdPayPal USD(PYUSD)$1.000.00%
  • MemeCoreMemeCore(M)$1.19-1.78%
  • OndoOndo(ONDO)$0.553.75%
  • tether-goldTether Gold(XAUT)$4,278.49-0.03%
  • okbOKB(OKB)$121.450.89%
  • Ripple USDRipple USD(RLUSD)$1.000.01%
  • Circle USYCCircle USYC(USYC)$1.140.00%
  • aaveAave(AAVE)$154.57-0.02%
  • Pump.funPump.fun(PUMP)$0.00507917.56%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.150.16%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Xiaomi Released MiMo-Audio, a 7B Speech Language Model Trained on 100M+ Hours with High-Fidelity Discrete Tokens

September 20, 2025
in AI & Technology
Reading Time: 8 mins read
A A
Xiaomi Released MiMo-Audio, a 7B Speech Language Model Trained on 100M+ Hours with High-Fidelity Discrete Tokens
ShareShareShareShareShare

Xiaomi’s MiMo team released MiMo-Audio, a 7-billion-parameter audio-language model that runs a single next-token objective over interleaved text and discretized speech, scaling pretraining beyond 100 million hours of audio.

What’s actually new?

Instead of relying on task-specific heads or lossy acoustic tokens, MiMo-Audio uses a bespoke RVQ (residual vector quantization) tokenizer that targets both semantic fidelity and high-quality reconstruction. The tokenizer runs at 25 Hz and outputs 8 RVQ layers (≈200 tokens/s), giving the LM access to “lossless” speech features it can model autoregressively alongside text.

YOU MAY ALSO LIKE

Bill Gates Says It’s ‘Completely Irresponsible’ For AI To Not Have Safeguards

Why The iPhone Duo Could Be Beneficial For Samsung’s Galaxy Z Fold 8

Architecture: patch encoder → 7B LLM → patch decoder

To handle the audio/text rate mismatch, the system packs four timesteps per patch for LM consumption (downsampling 25 Hz → 6.25 Hz), then reconstructs full-rate RVQ streams with a causal patch decoder. A delayed multi-layer RVQ generation scheme staggers predictions per codebook to stabilize synthesis and respect inter-layer dependencies. All three parts—patch encoder, MiMo-7B backbone, and patch decoder—are trained under a single next-token objective.

https://xiaomimimo.github.io/MiMo-Audio-Demo/

Scale is the algorithm

Training proceeds in two big phases: (1) an “understanding” stage that optimizes text-token loss over interleaved speech-text corpora, and (2) a joint “understanding + generation” stage that turns on audio losses for speech continuation, S2T/T2S tasks, and instruction-style data. The report emphasizes a compute/data threshold where few-shot behavior appears to “switch on,” echoing emergence curves seen in large text-only LMs.

Benchmarks: speech intelligence and general audio

MiMo-Audio is evaluated on speech-reasoning suites (e.g., SpeechMMLU) and broad audio understanding benchmarks (e.g., MMAU), reporting strong scores across speech, sound, and music and a reduced “modality gap” between text-only and speech-in/speech-out settings. Xiaomi also releases MiMo-Audio-Eval, a public toolkit to reproduce these results. Listen-and-respond demos (speech continuation, voice/emotion conversion, denoising, and speech translation) are available online.

https://xiaomimimo.github.io/MiMo-Audio-Demo/

Why this is important?

The approach is intentionally simple—no multi-head task tower, no bespoke ASR/TTS objectives at pretraining time—just GPT-style next-token prediction over lossless audio tokens plus text. The key engineering ideas are (i) a tokenizer the LM can actually use without throwing away prosody and speaker identity; (ii) patchification to keep sequence lengths manageable; and (iii) delayed RVQ decoding to preserve quality at generation time. For teams building spoken agents, those design choices translate into few-shot speech-to-speech editing and robust speech continuation with minimal task-specific finetuning.

6 Technical Takeaways:

  1. High-Fidelity Tokenization
    MiMo-Audio uses a custom RVQ tokenizer operating at 25 Hz with 8 active codebooks, ensuring speech tokens preserve prosody, timbre, and speaker identity while keeping them LM-friendly.
  2. Patchified Sequence Modeling
    The model reduces sequence length by grouping 4 timesteps into one patch (25 Hz → 6.25 Hz), letting the 7B LLM handle long speech efficiently without discarding detail.
  3. Unified Next-Token Objective
    Rather than separate heads for ASR, TTS, or dialogue, MiMo-Audio trains under a single next-token prediction loss across interleaved text and audio, simplifying architecture while supporting multi-task generalization.
  4. Emergent Few-Shot Abilities
    Few-shot behaviors such as speech continuation, voice conversion, emotion transfer, and speech translation emerge once training surpasses a large-scale data threshold (~100M hours, trillions of tokens).
  5. Benchmark Leadership
    MiMo-Audio sets state-of-the-art scores on SpeechMMLU (S2S 69.1, T2S 71.5) and MMAU (66.0 overall), while minimizing the text-to-speech modality gap to just 3.4 points.
  6. Open Ecosystem Release
    Xiaomi provides the tokenizer, 7B checkpoints (base and instruct), MiMo-Audio-Eval toolkit, and public demos, enabling researchers and developers to test and extend speech-to-speech intelligence in open-source pipelines.

Summary

MiMo-Audio demonstrates that high-fidelity, RVQ-based “lossless” tokenization combined with patchified next-token pretraining at scale is sufficient to unlock few-shot speech intelligence without task-specific heads. The 7B stack—tokenizer → patch encoder → LLM → patch decoder—bridges the audio/text rate gap (25→6.25 Hz) and preserves prosody and speaker identity via delayed multi-layer RVQ decoding. Empirically, the model narrows the text↔speech modality gap, generalizes across speech/sound/music benchmarks, and supports in-context S2S editing and continuation.


Check out the Paper, Technical details and GitHub Page. Feel free to check out our GitHub Page for Tutorials, Codes and Notebooks. Also, feel free to follow us on Twitter and don’t forget to join our 100k+ ML SubReddit and Subscribe to our Newsletter.


Michal Sutter is a data science professional with a Master of Science in Data Science from the University of Padova. With a solid foundation in statistical analysis, machine learning, and data engineering, Michal excels at transforming complex datasets into actionable insights.

🔥[Recommended Read] NVIDIA AI Open-Sources ViPE (Video Pose Engine): A Powerful and Versatile 3D Video Annotation Tool for Spatial AI

Credit: Source link

ShareTweetSendSharePin

Related Posts

Bill Gates Says It’s ‘Completely Irresponsible’ For AI To Not Have Safeguards
AI & Technology

Bill Gates Says It’s ‘Completely Irresponsible’ For AI To Not Have Safeguards

September 27, 2026
Why The iPhone Duo Could Be Beneficial For Samsung’s Galaxy Z Fold 8
AI & Technology

Why The iPhone Duo Could Be Beneficial For Samsung’s Galaxy Z Fold 8

September 27, 2026
How To Improve Your Router’s Security In 10 Minutes
AI & Technology

How To Improve Your Router’s Security In 10 Minutes

September 27, 2026
Humanoid Robots Are Getting Even Creepier (This One Can Cry On Command)
AI & Technology

Humanoid Robots Are Getting Even Creepier (This One Can Cry On Command)

September 27, 2026
Next Post
FCC Under Pressure After Disney Pulls Jimmy Kimmel’s ABC Show Off the Air

FCC Under Pressure After Disney Pulls Jimmy Kimmel's ABC Show Off the Air

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Why Your iPhone Hates Printer Cables

Why Your iPhone Hates Printer Cables

September 26, 2026
Orange: The Market Is Underestimating This European Telecom Giant (Rating Upgrade) (ORANY)

Orange: The Market Is Underestimating This European Telecom Giant (Rating Upgrade) (ORANY)

September 23, 2026
At Least 4 Dead After Explosion in Athens Tourist District – The New York Times

At Least 4 Dead After Explosion in Athens Tourist District – The New York Times

September 26, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!