Meta's Muse Voice Transcribe Posts 3.1% WER and $0.18 Hourly Price

Meta's Muse Voice Transcribe Posts 3.1% WER and $0.18 Hourly Price

N
News Editor
2026-09-02 09:43:40
Meta Superintelligence Labs has released Muse Voice Transcribe, a single model that combines real-time speech-to-text, speaker diarization and endpoint detection. The system can process audio longer than one hour and distinguish more than 20 speakers. In Artificial Analysis's English real-time transcription evaluation, Muse logged a word error rate of 3.1% (lower is better), ahead of GPT Live Transcribe at 3.9% and Gemini 3.5 Transcribe Live at 4.0%. Muse returns its final text about 0.16 seconds after a speaker stops. Rather than applying one fixed wait time to every word, the model reads 80ms of audio at a time and chooses whether to keep listening or start outputting. Easy words can be written out quickly, while harder ones get more context before being finalized. Meta used reinforcement learning to optimize accuracy and latency together. Muse was trained on more than 70 languages, 25 of which have been prioritized for validation, and it can recognize sentences that switch between Chinese and English. It is already integrated into Meta AI for Mac and Muse Code, and it is also offered through the Meta Model API. Pricing is $3 per 1,000 minutes, equal to $0.18 per hour.

Meta Superintelligence Labs has released Muse Voice Transcribe, a speech model that combines real-time transcription, speaker diarization and endpoint detection in one system. It can process audio streams longer than an hour and tell apart more than 20 voices in a single conversation.

In Artificial Analysis's English real-time transcription benchmark, Muse Voice Transcribe posted a word error rate of 3.1% (lower is better). GPT Live Transcribe came in at 3.9%, while Gemini 3.5 Transcribe Live finished at 4.0%.

Latency logic: decide every 80ms whether to keep listening

Muse returns final text about 0.16 seconds after the speaker stops. It does not assign one fixed wait time to every word. Instead, the model reads audio in 80ms slices and decides, slice by slice, whether to keep listening or begin writing. Simple words can be produced quickly; harder words get extra context before the system commits. Meta trained Muse with reinforcement learning, optimizing accuracy and latency at the same time.

Languages, availability and pricing

The model was trained on more than 70 languages, with 25 prioritized for focused validation. It can also recognize sentences that switch between Chinese and English mid-sentence.

Muse Voice Transcribe is now integrated into Meta AI for Mac and Muse Code, and it is available through the Meta Model API. Pricing is $3 per 1,000 minutes, which works out to $0.18 per hour.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
100

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.