Google has added a new transcription model to Gemini Audio, launching Gemini 3.5 Transcribe with support for more than 85 languages, automatic filler-word removal, and speaker separation for up to three people.
Google says the model can detect specialized terminology in conversations without requiring users to set the language in advance. It also supports custom vocabulary lists, letting users preload company names, product names, and unusual spellings so the transcript reflects those terms directly instead of forcing manual corrections later.
Up to three speakers in prerecorded audio
For prerecorded audio files, Gemini 3.5 Transcribe can identify as many as three speakers and attach word-level timestamps, giving users a cleaner base for editing and transcript search.
Google’s published accuracy figures
Transcription quality is commonly measured by word error rate, or WER, which compares machine-generated text with human transcription. Lower numbers indicate higher accuracy.
In Google’s published figures, the model records an average WER of 4.0% in streaming mode, where text appears as speech is being spoken, and 2.6% in non-streaming mode, where output is generated after the full segment is complete. On the multilingual FLEURS benchmark, the company lists 5.50% for streaming and 5.04% for non-streaming.
Compared with Chirp 3, Google says final transcription latency improved by 70%, cutting the wait between the end of speech and the appearance of a completed transcript.
APIs and rollout
Google is offering live transcription through the Live API under the model name gemini-3.5-transcribe-live. Prerecorded audio runs through the Interactions API under gemini-3.5-transcribe. The two tracks are positioned for different use cases.
The update is available starting today in English on the macOS Gemini app. Android’s Rambler dictation feature is opening in selected countries and languages. Developers can test the model during the Gemini API public preview through AI Studio and Antigravity, while Chrome support has not launched yet.
Fourth place on an external ranking
Google’s “major progress” claim is relative to its own previous model. In the broader transcription market, Gemini 3.5 Transcribe is neither the most accurate nor the lowest-priced option in the comparison cited by the source article.
Artificial Analysis, which publishes the AA-WER ranking, compares transcription models using the same yardstick for word error rate and price. On that list, Gemini 3.5 Transcribe posts a 2.6% WER at a price of $5 per 1,000 minutes.
Models ahead of it include ElevenLabs’ Scribe v2, which shows a 2.2% WER at $3.67 per 1,000 minutes, Microsoft Azure’s MAI-Transcribe-1.5 at 2.4%, and Smallest.ai’s Pulse Pro, also at 2.4%.
Further down the same ranking, AssemblyAI’s Universal-3 Pro posts 3.1% at $3.50 per 1,000 minutes, OpenAI’s GPT-4o Transcribe comes in at 4.0% and $6, and Deepgram’s Nova-3 records 5.2% at $4.30.
That comparison shows a crowded transcription field, with ElevenLabs, Microsoft, AssemblyAI, Deepgram, and OpenAI all competing on the same table across price and accuracy. On that measure, Google is still chasing the leaders rather than pulling away from them.
One factor working in Google’s favor is distribution. The source article notes that Scribe v2 depends on developers choosing to integrate its API, while Gemini 3.5 Transcribe can be pushed directly into the Gemini app, Android dictation features, and eventually Chrome, putting it in front of a much broader base of mainstream users.

