Microsoft AI on October 1, 2026 launched MAI-Transcribe-2-Streaming, its first streaming transcription model, alongside two new text-to-speech models, MAI-Voice-2.1 and MAI-Voice-2.1-Flash, with all three available through Microsoft Foundry.
MAI-Transcribe-2-Streaming and the Artificial Analysis Benchmark
Microsoft describes MAI-Transcribe-2-Streaming as delivering low-latency, real-time transcripts in 60 languages with automatic, continuous language detection. The company said the model ranks No. 1 for accuracy for both final and partial transcripts on Artificial Analysis, and that it sits on the Pareto frontier of the benchmark’s accuracy-versus-latency evaluation, meaning higher accuracy does not require a heavy latency tradeoff. The leaderboard chart in the post, citing the Artificial Analysis streaming leaderboard dated September 28, 2026, shows the model at a 2.5 percent final word-error rate, a 2.8 percent first-partial rate, and 0.13 seconds to final transcription.
Rather than waiting for a speaker to finish before returning text, the model produces its first hypotheses, known as partials, in just over 100 milliseconds of receiving audio, then revises them as more context arrives before committing a stable transcript. Microsoft said this allows voice-enabled applications to act on speech before the speaker finishes: voice agents can start reasoning or calling tools mid-sentence, and live transcripts can appear as people talk. For real-time dictation and subtitling, the company said its internal evaluations show words appearing in the transcript twice as fast as with its closest competitor.
Artificial Analysis states that its AA-WER Streaming index measures transcription accuracy for models where audio streams in real time, chunk by chunk, across roughly eight hours of audio from three datasets: AA-AgentTalk at 50 percent, VoxPopuli at 25 percent, and Earnings22 at 25 percent. The datasets cover real-world speech with diverse accents, domain-specific language, and challenging acoustic conditions, and the benchmark’s Time to Final and Time to First Partial measurements both start at the end of speech detected by the SileroVAD voice-activity detector.
MAI-Transcribe-2-Streaming is available at an introductory price of $0.54 per hour of audio through the end of the year. The model extends Microsoft’s MAI audio line, which already includes MAI-Transcribe-2, the earlier non-streaming speech recognition model the company billed as the fastest, most accurate and cheapest in the world.
MAI-Voice-2.1 and MAI-Voice-2.1-Flash
MAI-Voice-2.1 supports 23 languages and 26 locales, and Microsoft said a single voice can use all of them with a native accent, keeping the same speaker identity when switching languages. A tutoring app, in the company’s example, can switch languages mid-lesson without swapping teachers, and a multilingual assistant can reply in whatever language it is addressed in while still sounding like the same voice. The model is priced at $22 per 1M characters.
MAI-Voice-2.1-Flash supports the same languages and cross-language speakers but is built for high-volume, latency-sensitive workloads. It can generate up to 45 seconds of audio with an end-to-end latency of 150 milliseconds, and Microsoft said it delivers 55 percent faster model inference and is roughly 60 percent cheaper than comparable models. It is priced at $15 per 1M characters.
Both voice models support voice cloning across all supported languages using a few seconds of reference audio, with built-in consent guardrails that Microsoft said prevent misuse. In a 4,000-listener Turing test combining the two new voice models, 50.3 percent of listeners rated MAI-Voice as equally or more human-like than human recordings, Microsoft said.
Microsoft framed pairing MAI-Transcribe-2-Streaming with MAI-Voice-2.1-Flash as buying back time on both ends of a voice-agent loop, the sequence of hearing, understanding, deciding, and speaking within the window where a human still experiences the interaction as a conversation. Listed developer use cases include customer service agents that transcribe requests as they are spoken and respond in natural speech, multilingual assistants that detect the spoken language and reply in any of the 23 supported MAI-Voice languages, and interactive learning and media applications using distinct speakers for tutoring, role-play, simulations, narration, and conversational content.
Availability and the Chatter Demo
MAI-Voice-2.1 and MAI-Voice-2.1-Flash are available through OpenRouter. All three models are available through Microsoft Foundry, the MAI Playground, Vercel, and Azure Voice Live, with LiveKit listed as coming soon.
To show the models working together in a live agent, Microsoft built Chatter, a new demo in the MAI Playground that lets users talk to a voice assistant powered by the transcription and voice models.
Credit: Source link


























