
Google launches Gemini 3.5 Transcribe, its new model for multilingual speech recognition
Google has introduced Gemini 3.5 Transcribe, a new AI audio model for speech recognition and transcription that the company says is more accurate than previous versions and can automatically detect more than 85 languages. The model can turn conversational or unstructured speech into formatted text, remove filler words, and let users edit transcripts with voice commands.
It also supports custom vocabulary and unusual spellings, and Google says it is better at capturing alphanumeric strings such as order numbers and postal codes. For prerecorded audio, Gemini 3.5 Transcribe can identify up to three speakers and provide word-level timestamps, making it useful for podcasts, interviews, meetings, and other multi-speaker recordings. The model already powers Gboard Rambler on Android devices including the Pixel 11 series, as well as transcription features in the Gemini app for macOS.
Google plans to integrate Gemini 3.5 Transcribe into Chrome so users can dictate text in web fields, including replies, posts, forms, and Gemini prompts. The model is also available in Google Antigravity and is coming to Search Live, Gemini Live, Docs, Keep, and Gmail, while developers will be able to access it through APIs.


