• bitcoinBitcoin(BTC)$86,747.001.57%
  • ethereumEthereum(ETH)$2,768.131.29%
  • tetherTether(USDT)$1.000.01%
  • binancecoinBNB(BNB)$791.840.68%
  • rippleXRP(XRP)$1.627.46%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$118.992.03%
  • tronTRON(TRX)$0.343873-1.47%
  • zcashZcash(ZEC)$1,619.147.01%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.031.77%
  • HyperliquidHyperliquid(HYPE)$97.533.64%
  • dogecoinDogecoin(DOGE)$0.1024593.35%
  • moneroMonero(XMR)$574.14-0.50%
  • whitebitWhiteBIT Coin(WBT)$87.201.51%
  • chainlinkChainlink(LINK)$13.061.05%
  • cardanoCardano(ADA)$0.2584465.82%
  • USDSUSDS(USDS)$1.00-0.01%
  • RainRain(RAIN)$0.013136-4.09%
  • leo-tokenLEO Token(LEO)$8.980.26%
  • stellarStellar(XLM)$0.2208294.49%
  • bitcoin-cashBitcoin Cash(BCH)$339.1928.12%
  • uniswapUniswap(UNI)$10.5617.56%
  • nearNEAR Protocol(NEAR)$4.472.80%
  • litecoinLitecoin(LTC)$64.285.40%
  • avalanche-2Avalanche(AVAX)$11.225.39%
  • Ethena USDeEthena USDe(USDE)$1.000.03%
  • daiDai(DAI)$1.00-0.01%
  • CantonCanton(CC)$0.116266-1.71%
  • USD1USD1(USD1)$1.000.00%
  • hedera-hashgraphHedera(HBAR)$0.0996208.05%
  • suiSui(SUI)$1.031.85%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.472.23%
  • shiba-inuShiba Inu(SHIB)$0.0000062.86%
  • BittensorBittensor(TAO)$315.53-0.85%
  • crypto-com-chainCronos(CRO)$0.0682243.62%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • MemeCoreMemeCore(M)$1.31-3.81%
  • paypal-usdPayPal USD(PYUSD)$1.000.00%
  • tether-goldTether Gold(XAUT)$4,333.660.23%
  • okbOKB(OKB)$125.463.40%
  • BitwayBitway(BTW)$0.9315.37%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • Ripple USDRipple USD(RLUSD)$1.00-0.01%
  • aaveAave(AAVE)$150.995.16%
  • mantleMantle(MNT)$0.698.25%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.150.05%
  • EthenaEthena(ENA)$0.2180740.16%
  • OndoOndo(ONDO)$0.4413721.76%
  • pepePepe(PEPE)$0.000005-1.20%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Microsoft Releases VibeVoice-ASR: A Unified Speech-to-Text Model Designed to Handle 60-Minute Long-Form Audio in a Single Pass

January 22, 2026
in AI & Technology
Reading Time: 6 mins read
A A
Microsoft Releases VibeVoice-ASR: A Unified Speech-to-Text Model Designed to Handle 60-Minute Long-Form Audio in a Single Pass
ShareShareShareShareShare

Microsoft has released VibeVoice-ASR as part of the VibeVoice family of open source frontier voice AI models. VibeVoice-ASR is described as a unified speech-to-text model that can handle 60-minute long-form audio in a single pass and output structured transcriptions that encode Who, When, and What, with support for Customized Hotwords.

VibeVoice sits in a single repository that hosts Text-to-Speech, real time TTS, and Automatic Speech Recognition models under an MIT license. VibeVoice uses continuous speech tokenizers that run at 7.5 Hz and a next-token diffusion framework where a Large Language Model reasons over text and dialogue and a diffusion head generates acoustic detail. This framework is mainly documented for TTS, but it defines the overall design context in which VibeVoice-ASR lives.

YOU MAY ALSO LIKE

OpenAI Releases GPT-6 Sol and Luna: 50% Cheaper API Pricing and Benchmarks

The Pros And Cons Of Using A Password Manager Over An Authenticator App

https://huggingface.co/microsoft/VibeVoice-ASR

Long form ASR with a single global context

Unlike conventional ASR (Automatic Speech Recognition) systems that first cut audio into short segments and then run diarization and alignment as separate components, VibeVoice-ASR is designed to accept up to 60 minutes of continuous audio input within a 64K token length budget. The model keeps one global representation of the full session. This means the model can maintain speaker identity and topic context across the entire hour instead of resetting every few seconds.

60-minute Single-Pass Processing

The first key feature is that many conventional ASR systems process long audio by cutting it into short segments, which can lose global context. VibeVoice-ASR instead takes up to 60 minutes of continuous audio within a 64K token window so it can maintain consistent speaker tracking and semantic context across the entire recording.

This is important for tasks like meeting transcription, lectures, and long support calls. A single pass over the complete sequence simplifies the pipeline. There is no need to implement custom logic to merge partial hypotheses or repair speaker labels at boundaries between audio chunks.

Customized Hotwords for domain accuracy

Customized Hotwords are the second key feature. Users can provide hotwords such as product names, organization names, technical terms, or background context. The model uses these hotwords to guide the recognition process.

This allows you to bias decoding toward the correct spelling and pronunciation for domain specific tokens without retraining the model. For example, a dev-user can pass internal project names or customer specific terms at inference time. This is useful when deploying the same base model across several products that share similar acoustic conditions but very different vocabularies.

Microsoft also ships a finetuning-asr directory with LoRA based fine tuning scripts for VibeVoice-ASR. Together, hotwords and LoRA fine tuning give a path for both light weight adaptation and deeper domain specialization.

Rich Transcription, diarization, and timing

The third feature is Rich Transcription with Who, When, and What. The model jointly performs ASR, diarization, and timestamping, and returns a structured output that indicates who said what and when.

See below the three evaluation figures named DER, cpWER, and tcpWER.

https://huggingface.co/microsoft/VibeVoice-ASR
  • DER is Diarization Error Rate, it measures how well the model assigns speech segments to the correct speaker
  • cpWER and tcpWER are word error rate metrics computed under conversational settings

These graphs summarize how well the model performs on multi speaker long form data, which is the primary target setting for this ASR system.

The structured output format is well suited for downstream processing like speaker specific summarization, action item extraction, or analytics dashboards. Since segments, speakers, and timestamps already come from a single model, downstream code can treat the transcript as a time aligned event log.

Key Takeaways

  • VibeVoice-ASR is a unified speech to text model that handles 60 minute long form audio in a single pass within a 64K token context.
  • The model jointly performs ASR, diarization, and timestamping so it outputs structured transcripts that encode Who, When, and What in a single inference step.
  • Customized Hotwords let users inject domain specific terms such as product names or technical jargon to improve recognition accuracy without retraining the model.
  • Evaluation with DER, cpWER, and tcpWER focuses on multi speaker conversational scenarios which aligns the model with meetings, lectures, and long calls.
  • VibeVoice-ASR is released in the VibeVoice open source stack under MIT license with official weights, fine tuning scripts, and an online Playground for experimentation.

Check out the Model Weights, Repo and Playground. Also, feel free to follow us on Twitter and don’t forget to join our 100k+ ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well.

The post Microsoft Releases VibeVoice-ASR: A Unified Speech-to-Text Model Designed to Handle 60-Minute Long-Form Audio in a Single Pass appeared first on MarkTechPost.

Credit: Source link

ShareTweetSendSharePin

Related Posts

OpenAI Releases GPT-6 Sol and Luna: 50% Cheaper API Pricing and Benchmarks
AI & Technology

OpenAI Releases GPT-6 Sol and Luna: 50% Cheaper API Pricing and Benchmarks

September 23, 2026
The Pros And Cons Of Using A Password Manager Over An Authenticator App
AI & Technology

The Pros And Cons Of Using A Password Manager Over An Authenticator App

September 23, 2026
How To Hide Or Replace The Audio Button In iMessages
AI & Technology

How To Hide Or Replace The Audio Button In iMessages

September 22, 2026
Improve Your Apple CarPlay Experience By Doing These Simple Things
AI & Technology

Improve Your Apple CarPlay Experience By Doing These Simple Things

September 22, 2026
Next Post
America’s best and worst airlines for 2025 revealed in new rankings

America’s best and worst airlines for 2025 revealed in new rankings

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Stanford Researchers Release Paper2Agent: Turning Research Papers Into AI Agents That Reproduce Results and Run on New Data

Stanford Researchers Release Paper2Agent: Turning Research Papers Into AI Agents That Reproduce Results and Run on New Data

September 16, 2026
Ukrainian troops innovate as drone warfare intensifies

Ukrainian troops innovate as drone warfare intensifies

September 17, 2026
Trump says he’s called for a ban on diesel exports – The Washington Post

Trump says he’s called for a ban on diesel exports – The Washington Post

September 23, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!