Microsoft has released three audio models at once — MAI-Transcribe-2-Streaming, MAI-Voice-2.1 and MAI-Voice-2.1-Flash — covering both real-time speech transcription and text-to-speech generation for low-latency voice agents.
One launch covers transcription and speech generation
MAI-Transcribe-2-Streaming handles live speech-to-text and can keep outputting text before a speaker has finished talking. It supports 60 languages and automatic language detection. MAI-Voice-2.1 and MAI-Voice-2.1-Flash are designed to turn text into speech.
Accuracy and response-time metrics
According to Artificial Analysis’ streaming speech transcription ranking, MAI-Transcribe-2-Streaming ranked first in both final transcription accuracy and first partial transcription accuracy. The model posted a 2.5% word error rate. For comparison, Grok Voice Transcribe 2.0 recorded a 2.7% word error rate, and its final text was returned about 0.49 seconds after the end of speech was detected. MAI-Transcribe-2-Streaming returned final text about 0.13 seconds after speech end detection.
Flash version targets lower latency
MAI-Voice-2.1 supports 23 languages and lets the same voice switch across different languages. The Flash version is tuned for lower latency, with model inference at about 45ms and pricing at $15 per million characters. The standard version runs at about 550ms and costs $22 per million characters.
Both text-to-speech models also support voice matching from a short reference audio clip.
Built to cut waiting time in voice agents
The model stack is aimed at one of the biggest friction points in voice agents: waiting. The system can start transcription and inference while the user is still speaking, then use the Flash model to generate a spoken reply quickly instead of waiting for a full sentence to finish before processing begins.

