Google rolls out Gemini 3.5 Transcribe with support for 85+ languages and three-speaker diarization
Google has introduced Gemini 3.5 Transcribe for Gemini Audio, adding a new speech-to-text model that can automatically recognize more than 85 languages, identify up to three speakers in prerecorded audio, and remove filler words such as “um” and “ah” during transcription. The company also says the model can pick up specialized terminology without requiring users to preset the language, and it supports custom vocabulary lists for company names, product names, and unusual spellings. According to Google’s published figures, Gemini 3.5 Transcribe posts an average word error rate of 4.0% in streaming mode and 2.6% in non-streaming mode, while scoring 5.50% and 5.04% respectively on the multilingual FLEURS benchmark. Google says final transcription latency is down 70% from Chirp 3. The model is now live in English on the macOS Gemini app, with limited availability for Android’s Rambler dictation in some countries and languages, while developers can access it through public preview in Gemini API tools. Still, third-party rankings from Artificial Analysis place the model fourth on accuracy, behind offerings from ElevenLabs, Microsoft Azure, and Smallest.ai.








