Google rolls out Gemini 3.5 Transcribe with support for 85+ languages and three-speaker diarization

Google rolls out Gemini 3.5 Transcribe with support for 85+ languages and three-speaker diarization

N
News Editor
2026-08-27 03:16:25
Google has introduced Gemini 3.5 Transcribe for Gemini Audio, adding a new speech-to-text model that can automatically recognize more than 85 languages, identify up to three speakers in prerecorded audio, and remove filler words such as “um” and “ah” during transcription. The company also says the model can pick up specialized terminology without requiring users to preset the language, and it supports custom vocabulary lists for company names, product names, and unusual spellings. According to Google’s published figures, Gemini 3.5 Transcribe posts an average word error rate of 4.0% in streaming mode and 2.6% in non-streaming mode, while scoring 5.50% and 5.04% respectively on the multilingual FLEURS benchmark. Google says final transcription latency is down 70% from Chirp 3. The model is now live in English on the macOS Gemini app, with limited availability for Android’s Rambler dictation in some countries and languages, while developers can access it through public preview in Gemini API tools. Still, third-party rankings from Artificial Analysis place the model fourth on accuracy, behind offerings from ElevenLabs, Microsoft Azure, and Smallest.ai.

Google has added a new transcription model to Gemini Audio, launching Gemini 3.5 Transcribe with support for more than 85 languages, automatic filler-word removal, and speaker separation for up to three people.

Google says the model can detect specialized terminology in conversations without requiring users to set the language in advance. It also supports custom vocabulary lists, letting users preload company names, product names, and unusual spellings so the transcript reflects those terms directly instead of forcing manual corrections later.

Up to three speakers in prerecorded audio

For prerecorded audio files, Gemini 3.5 Transcribe can identify as many as three speakers and attach word-level timestamps, giving users a cleaner base for editing and transcript search.

Google’s published accuracy figures

Transcription quality is commonly measured by word error rate, or WER, which compares machine-generated text with human transcription. Lower numbers indicate higher accuracy.

In Google’s published figures, the model records an average WER of 4.0% in streaming mode, where text appears as speech is being spoken, and 2.6% in non-streaming mode, where output is generated after the full segment is complete. On the multilingual FLEURS benchmark, the company lists 5.50% for streaming and 5.04% for non-streaming.

Compared with Chirp 3, Google says final transcription latency improved by 70%, cutting the wait between the end of speech and the appearance of a completed transcript.

APIs and rollout

Google is offering live transcription through the Live API under the model name gemini-3.5-transcribe-live. Prerecorded audio runs through the Interactions API under gemini-3.5-transcribe. The two tracks are positioned for different use cases.

The update is available starting today in English on the macOS Gemini app. Android’s Rambler dictation feature is opening in selected countries and languages. Developers can test the model during the Gemini API public preview through AI Studio and Antigravity, while Chrome support has not launched yet.

Fourth place on an external ranking

Google’s “major progress” claim is relative to its own previous model. In the broader transcription market, Gemini 3.5 Transcribe is neither the most accurate nor the lowest-priced option in the comparison cited by the source article.

Artificial Analysis, which publishes the AA-WER ranking, compares transcription models using the same yardstick for word error rate and price. On that list, Gemini 3.5 Transcribe posts a 2.6% WER at a price of $5 per 1,000 minutes.

Models ahead of it include ElevenLabs’ Scribe v2, which shows a 2.2% WER at $3.67 per 1,000 minutes, Microsoft Azure’s MAI-Transcribe-1.5 at 2.4%, and Smallest.ai’s Pulse Pro, also at 2.4%.

Further down the same ranking, AssemblyAI’s Universal-3 Pro posts 3.1% at $3.50 per 1,000 minutes, OpenAI’s GPT-4o Transcribe comes in at 4.0% and $6, and Deepgram’s Nova-3 records 5.2% at $4.30.

That comparison shows a crowded transcription field, with ElevenLabs, Microsoft, AssemblyAI, Deepgram, and OpenAI all competing on the same table across price and accuracy. On that measure, Google is still chasing the leaders rather than pulling away from them.

One factor working in Google’s favor is distribution. The source article notes that Scribe v2 depends on developers choosing to integrate its API, while Gemini 3.5 Transcribe can be pushed directly into the Gemini app, Android dictation features, and eventually Chrome, putting it in front of a much broader base of mainstream users.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
60

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.