Microsoft rolls out three audio models for voice agents, with speech generation latency as low as 45ms

Microsoft rolls out three audio models for voice agents, with speech generation latency as low as 45ms

N
News Editor
2026-10-03 02:32:34
Microsoft has introduced three audio models in one move: MAI-Transcribe-2-Streaming, MAI-Voice-2.1 and MAI-Voice-2.1-Flash. Together, they cover the full voice-agent pipeline, from real-time speech-to-text to text-to-speech generation, with a focus on low-latency use cases. MAI-Transcribe-2-Streaming is built for live transcription and can keep producing text before a speaker finishes a sentence. It supports 60 languages and automatic language detection. According to Artificial Analysis’ streaming speech transcription ranking, the model placed first in both final transcription accuracy and first partial transcription accuracy, posting a 2.5% word error rate. Its final text is returned about 0.13 seconds after speech end is detected, compared with 2.7% and 0.49 seconds for Grok Voice Transcribe 2.0. On the speech generation side, MAI-Voice-2.1 supports 23 languages and allows a single voice to switch across languages. The Flash version is tuned for speed, with inference at about 45ms and pricing of $15 per million characters, while the standard version runs at about 550ms and costs $22 per million characters. Both text-to-speech models can match a voice using a short reference audio sample.

Microsoft has released three audio models at once — MAI-Transcribe-2-Streaming, MAI-Voice-2.1 and MAI-Voice-2.1-Flash — covering both real-time speech transcription and text-to-speech generation for low-latency voice agents.

One launch covers transcription and speech generation

MAI-Transcribe-2-Streaming handles live speech-to-text and can keep outputting text before a speaker has finished talking. It supports 60 languages and automatic language detection. MAI-Voice-2.1 and MAI-Voice-2.1-Flash are designed to turn text into speech.

Accuracy and response-time metrics

According to Artificial Analysis’ streaming speech transcription ranking, MAI-Transcribe-2-Streaming ranked first in both final transcription accuracy and first partial transcription accuracy. The model posted a 2.5% word error rate. For comparison, Grok Voice Transcribe 2.0 recorded a 2.7% word error rate, and its final text was returned about 0.49 seconds after the end of speech was detected. MAI-Transcribe-2-Streaming returned final text about 0.13 seconds after speech end detection.

Flash version targets lower latency

MAI-Voice-2.1 supports 23 languages and lets the same voice switch across different languages. The Flash version is tuned for lower latency, with model inference at about 45ms and pricing at $15 per million characters. The standard version runs at about 550ms and costs $22 per million characters.

Both text-to-speech models also support voice matching from a short reference audio clip.

Built to cut waiting time in voice agents

The model stack is aimed at one of the biggest friction points in voice agents: waiting. The system can start transcription and inference while the user is still speaking, then use the Flash model to generate a spoken reply quickly instead of waiting for a full sentence to finish before processing begins.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
100

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.