• bitcoinBitcoin(BTC)$81,261.000.48%
  • ethereumEthereum(ETH)$2,634.630.94%
  • tetherTether(USDT)$1.00-0.01%
  • binancecoinBNB(BNB)$762.210.07%
  • rippleXRP(XRP)$1.411.13%
  • usd-coinUSDC(USDC)$1.00-0.01%
  • solanaSolana(SOL)$111.07-1.49%
  • tronTRON(TRX)$0.3397840.44%
  • zcashZcash(ZEC)$1,471.76-6.29%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.02-1.43%
  • HyperliquidHyperliquid(HYPE)$92.00-0.63%
  • dogecoinDogecoin(DOGE)$0.0879030.29%
  • moneroMonero(XMR)$546.08-3.12%
  • whitebitWhiteBIT Coin(WBT)$82.91-0.13%
  • RainRain(RAIN)$0.0137542.90%
  • USDSUSDS(USDS)$1.00-0.02%
  • chainlinkChainlink(LINK)$12.431.66%
  • cardanoCardano(ADA)$0.2278891.58%
  • leo-tokenLEO Token(LEO)$8.90-0.09%
  • stellarStellar(XLM)$0.1966732.12%
  • uniswapUniswap(UNI)$8.66-2.12%
  • bitcoin-cashBitcoin Cash(BCH)$256.440.61%
  • Ethena USDeEthena USDe(USDE)$1.00-0.02%
  • nearNEAR Protocol(NEAR)$3.58-4.14%
  • daiDai(DAI)$1.000.02%
  • litecoinLitecoin(LTC)$58.090.38%
  • avalanche-2Avalanche(AVAX)$10.1423.39%
  • USD1USD1(USD1)$1.00-0.02%
  • CantonCanton(CC)$0.109411-1.78%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.380.61%
  • hedera-hashgraphHedera(HBAR)$0.0815913.05%
  • suiSui(SUI)$0.865.84%
  • MemeCoreMemeCore(M)$1.5115.23%
  • Global DollarGlobal Dollar(USDG)$1.00-0.01%
  • shiba-inuShiba Inu(SHIB)$0.0000061.05%
  • BittensorBittensor(TAO)$264.066.68%
  • crypto-com-chainCronos(CRO)$0.059523-0.36%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.03%
  • tether-goldTether Gold(XAUT)$4,373.76-0.06%
  • Circle USYCCircle USYC(USYC)$1.140.00%
  • okbOKB(OKB)$117.751.17%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.15-0.29%
  • aaveAave(AAVE)$141.441.40%
  • AsterAster(ASTER)$0.76-0.68%
  • EthenaEthena(ENA)$0.20393020.90%
  • mantleMantle(MNT)$0.62-0.37%
  • OndoOndo(ONDO)$0.4205555.61%
  • Pump.funPump.fun(PUMP)$0.004248-0.54%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Microsoft Released VibeVoice-1.5B: An Open-Source Text-to-Speech Model that can Synthesize up to 90 Minutes of Speech with Four Distinct Speakers

August 25, 2025
in AI & Technology
Reading Time: 7 mins read
A A
Microsoft Released VibeVoice-1.5B: An Open-Source Text-to-Speech Model that can Synthesize up to 90 Minutes of Speech with Four Distinct Speakers
ShareShareShareShareShare

Microsoft’s latest open source release, VibeVoice-1.5B, redefines the boundaries of text-to-speech (TTS) technology—delivering expressive, long-form, multi-speaker generated audio that is MIT licensed, scalable, and highly flexible for research use. This model isn’t just another TTS engine; it’s a framework designed to generate up to 90 minutes of uninterrupted, natural-sounding audio, support simultaneous generation of up to four distinct speakers, and even handle cross-lingual and singing synthesis scenarios. With a streaming architecture and a larger 7B model announced for the near future, VibeVoice-1.5B positions itself as a major advance for AI-powered conversational audio, podcasting, and synthetic voice research.

Key Features

  • Massive Context and Multi-Speaker Support: VibeVoice-1.5B can synthesize up to 90 minutes of speech with up to four distinct speakers in a single session—far surpassing the typical 1-2 speaker limit of traditional TTS models.
  • Simultaneous Generation: The model isn’t just stitching together single-voice clips; it’s designed to support parallel audio streams for multiple speakers, mimicking natural conversation and turn-taking.
  • Cross-Lingual and Singing Synthesis: While primarily trained on English and Chinese, the model is capable of cross-lingual synthesis and can even generate singing—features rarely demonstrated in previous open source TTS models.
  • MIT License: Fully open source and commercially friendly, with a focus on research, transparency, and reproducibility.
  • Scalable for Streaming and Long-Form Audio: The architecture is designed for efficient long-duration synthesis and anticipates a forthcoming 7B streaming-capable model, further expanding possibilities for real-time and high-fidelity TTS.
  • Emotion and Expressiveness: The model is touted for its emotion control and natural expressiveness, making it suitable for applications like podcasts or conversational scenarios.
https://huggingface.co/microsoft/VibeVoice-1.5B

Architecture and Technical Deep Dive

VibeVoice’s foundation is a 1.5B-parameter LLM (Qwen2.5-1.5B) that integrates with two novel tokenizers—Acoustic and Semantic—both designed to operate at a low frame rate (7.5Hz) for computational efficiency and consistency across long sequences.

YOU MAY ALSO LIKE

Trump Proposes Renaming Artificial Intelligence, Announces AI Force – Unite.AI

SpaceX Targets September 28 For Starship’s First Orbital Flight

  • Acoustic Tokenizer: A σ-VAE variant with a mirrored encoder-decoder structure (each ~340M parameters), achieving 3200x downsampling from raw audio at 24kHz.
  • Semantic Tokenizer: Trained via an ASR proxy task, this encoder-only architecture mirrors the acoustic tokenizer’s design (minus the VAE components).
  • Diffusion Decoder Head: A lightweight (~123M parameter) conditional diffusion module predicts acoustic features, leveraging Classifier-Free Guidance (CFG) and DPM-Solver for perceptual quality.
  • Context Length Curriculum: Training starts at 4k tokens and scales up to 65k tokens—enabling the model to generate very long, coherent audio segments.
  • Sequence Modeling: The LLM understands dialogue flow for turn-taking, while the diffusion head generates fine-grained acoustic details—separating semantics and synthesis while preserving speaker identity over long durations.

Model Limitations and Responsible Use

  • English and Chinese Only: The model is trained solely on these languages; other languages may produce unintelligible or offensive outputs.
  • No Overlapping Speech: While it supports turn-taking, VibeVoice-1.5B does not model overlapping speech between speakers.
  • Speech-Only: The model does not generate background sounds, Foley, or music—audio output is strictly speech.
  • Legal and Ethical Risks: Microsoft explicitly prohibits use for voice impersonation, disinformation, or authentication bypass. Users must comply with laws and disclose AI-generated content.
  • Not for Professional Real-Time Applications: While efficient, this release is not optimized for low-latency, interactive, or live-streaming scenarios; that’s the target for the soon-to-come 7B variant.

Conclusion

Microsoft’s VibeVoice-1.5B is a breakthrough in open TTS: scalable, expressive, and multi-speaker, with a lightweight diffusion-based architecture that unlocks long-form, conversational audio synthesis for researchers and open source developers. While use is currently research-focused and limited to English/Chinese, the model’s capabilities—and the promise of upcoming versions—signal a paradigm shift in how AI can generate and interact with synthetic speech.

For technical teams, content creators, and AI enthusiasts, VibeVoice-1.5B is a must-explore tool for the next generation of synthetic voice applications—available now on Hugging Face and GitHub, with clear documentation and an open license. As the field pivots toward more expressive, interactive, and ethically transparent TTS, Microsoft’s latest offering is a landmark for open source AI speech synthesis.


FAQs

What makes VibeVoice-1.5B different from other text-to-speech models?

VibeVoice-1.5B can generate up to 90 minutes of expressive, multi-speaker audio (up to four speakers), supports cross-lingual and singing synthesis, and is fully open source under the MIT license—pushing the boundaries of long-form conversational AI audio generation

What hardware is recommended for running the model locally?

Community tests show that generating a multi-speaker dialog with the 1.5 B checkpoint consumes ≈ 7 GB of GPU VRAM, so an 8 GB consumer card (e.g., RTX 3060) is generally sufficient for inference.

Which languages and audio styles does the model support today?

VibeVoice-1.5B is trained only on English and Chinese and can perform cross-lingual narration (e.g., English prompt → Chinese speech) as well as basic singing synthesis. It produces speech only—no background sounds—and does not model overlapping speakers; turn-taking is sequential.


Check out the Technical Report, Model on Hugging Face and Codes. Feel free to check out our GitHub Page for Tutorials, Codes and Notebooks. Also, feel free to follow us on Twitter and don’t forget to join our 100k+ ML SubReddit and Subscribe to our Newsletter.


Asif Razzaq is the CEO of Marktechpost Media Inc.. As a visionary entrepreneur and engineer, Asif is committed to harnessing the potential of Artificial Intelligence for social good. His most recent endeavor is the launch of an Artificial Intelligence Media Platform, Marktechpost, which stands out for its in-depth coverage of machine learning and deep learning news that is both technically sound and easily understandable by a wide audience. The platform boasts of over 2 million monthly views, illustrating its popularity among audiences.

Credit: Source link

ShareTweetSendSharePin

Related Posts

Trump Proposes Renaming Artificial Intelligence, Announces AI Force – Unite.AI
AI & Technology

Trump Proposes Renaming Artificial Intelligence, Announces AI Force – Unite.AI

September 19, 2026
SpaceX Targets September 28 For Starship’s First Orbital Flight
AI & Technology

SpaceX Targets September 28 For Starship’s First Orbital Flight

September 19, 2026
Now Trump Says He’s Creating An AI Force
AI & Technology

Now Trump Says He’s Creating An AI Force

September 19, 2026
TypeSafe AI Releases Jev: A System One Model That Returns Typed, Calibrated Decisions Instead of Text
AI & Technology

TypeSafe AI Releases Jev: A System One Model That Returns Typed, Calibrated Decisions Instead of Text

September 19, 2026
Next Post
New memoir by CBRE’s Stephen Siegel an entertaining look at powerful NYC dealmaker’s career

New memoir by CBRE's Stephen Siegel an entertaining look at powerful NYC dealmaker's career

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
People's houses are collapsing into the ocean. FEMA gives them no other option – NPR

People's houses are collapsing into the ocean. FEMA gives them no other option – NPR

September 17, 2026
Meet the Press NOW — September 10

Meet the Press NOW — September 10

September 14, 2026
Will Lindsay Clancy be retried? How and when it could happen

Will Lindsay Clancy be retried? How and when it could happen

September 17, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!