Microsoft AI has introduced two new models, MAI-Transcribe-2-Streaming and MAI-TTS-2, aimed at improving capabilities for voice agents. According to a Techub News item citing The Decoder, the first model focuses on real-time speech-to-text transcription, while the second converts text into natural-sounding speech. Both were tuned for voice-agent use cases rather than presented as general-purpose releases. The announcement centers on two core functions that sit at the heart of spoken AI systems: converting live audio into text with low latency, and generating spoken responses from written input. In the brief release, Microsoft AI framed the pair as tools designed to strengthen speech-based agent experiences. No additional technical specifications, launch regions, pricing details, or deployment timelines were disclosed in the source item.
Microsoft AI has released two new models, MAI-Transcribe-2-Streaming and MAI-TTS-2, according to a Techub News brief that cited The Decoder. The models are intended to improve voice agent capabilities.
MAI-Transcribe-2-Streaming is designed for real-time speech-to-text transcription. MAI-TTS-2 converts text into natural speech. Both models were optimized for voice agent application scenarios.
This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan. Disclaimer:
The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.
Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.