Quick Facts

  • Microsoft's MAI-Transcribe-2-Streaming ranks first among 38 models on the Artificial Analysis streaming benchmark with a 2.5% word error rate and 0.13-second final transcription time.
  • The streaming transcription model is priced at $0.54 per hour of audio through the end of 2026, while the new MAI-Voice-2.1-Flash text-to-speech model costs $15 per million characters.
  • All three models are available now through Microsoft Foundry and the MAI Playground.

Microsoft expanded its MAI model family on Oct. 1 with three new voice AI models aimed at developers building real-time conversational applications. The release includes MAI-Transcribe-2-Streaming, the company's first streaming transcription model, along with two text-to-speech models: MAI-Voice-2.1 and MAI-Voice-2.1-Flash.

MAI-Transcribe-2-Streaming accepts audio via a WebSocket and begins producing transcript hypotheses in just over 100 milliseconds. It revises those partial results as more audio arrives and confirms a final transcript 0.13 seconds after the speaker finishes. The model lets voice agents begin reasoning or calling tools before a user stops speaking.

According to SiliconAngle, the model holds the top spot among 38 models on the Artificial Analysis AA-WER Streaming benchmark, posting a 2.5% final word error rate. Its nearest competitor, Grok Voice Transcribe 2.0, scored 2.73%. Microsoft said the model sits on the Pareto frontier of the benchmark's accuracy-versus-latency evaluation, meaning gains in accuracy do not require added latency.

Microsoft AI chief Mustafa Suleyman called it "the most accurate real-time transcription model in the world" and said it is 55% faster and 60% cheaper than ElevenLabs. For context, Artificial Analysis lists ElevenLabs Scribe v2 Realtime and Deepgram Flux both at $6.50 per 1,000 minutes. Microsoft's introductory rate of $0.54 per hour converts to roughly $9 per 1,000 minutes.

The two new text-to-speech models serve distinct use cases. MAI-Voice-2.1 is Microsoft's highest-fidelity option, covering 23 languages with support for instant voice cloning and long-form audio generation. A developer building a tutoring app, for example, could move between English and Mandarin explanations while keeping a consistent voice throughout.

MAI-Voice-2.1-Flash is built for high-volume, latency-sensitive workloads. It generates 45 seconds of audio in 150 milliseconds and delivers 55% faster model inference compared to its predecessor. At $15 per million characters, it is roughly 32% cheaper than MAI-Voice-2.1, which costs $22 per million characters. Both models support voice cloning from a few seconds of reference audio and include consent guardrails against misuse.

The streaming model is priced at a significant premium over Microsoft's existing batch transcription offering. Last month, the company released MAI-Transcribe-2 in batch mode at $0.10 per audio hour. The streaming variant costs more than five times that, reflecting the additional compute required to process and return provisional results continuously.

Microsoft internally claims the streaming model displays words twice as fast as its closest competitor for dictation and subtitle use cases, though the company did not name the specific competitor in that comparison.

All three models are accessible through Microsoft Foundry and the MAI Playground. The launch adds voice input and output to a growing set of MAI developer tools and positions Microsoft to compete directly with dedicated voice AI vendors such as ElevenLabs and Deepgram in the enterprise developer market.

Read more: Microsoft targets ultra-realistic voice agents with its first streaming transcription model

This article was written by an AI agent. Spotted an error? Send a correction and we will fix it.