Google Gemini 3.5 Transcribe Launches With 2.6% WER and 85 Languages
Google launches Gemini 3.5 Transcribe with sub-second streaming, 85-language support, smart speech cleanup and a 2.6% non-streaming error rate for developers.
Summary
On Aug. 26, 2026, Google introduced Gemini 3.5 Transcribe, a speech-to-text model that converts raw audio into polished, formatted text for voice agents, captions and post-call analytics. gemini-3.5-transcribe-live delivers bidirectional Live API streaming with sub-second latency; gemini-3.5-transcribe uses the Interactions API for recordings, speaker attribution and word-level timestamps. Artificial Analysis measured average WER of 4.0% streaming and 2.6% non-streaming, plus 70% faster final transcription than Chirp 3. FLEURS WER reached 5.50% streaming and 5.04% non-streaming.
It detects more than 85 languages, regional accents and dialects, handles live language switching, noise, fillers, self-corrections, custom vocabulary, postal codes and order IDs, and attributes up to three speakers; support beyond three remains experimental. Function calls can send image generation and file analysis to other Gemini models, currently in Gemini for macOS. Gboard’s Rambler cleans dictation and supports voice edits; Antigravity can use permitted screen context and chat history; AI Studio Build supports voice coding; Gemini for macOS combines voice commands with screen context; Chrome talk-to-type is coming.
Developers can access public previews through the Gemini API in Google AI Studio and Google Antigravity; enterprises through Gemini Enterprise Agent Platform, with Gemini Enterprise for Customer Experience coming soon. Consumer access covers Gemini for macOS in English and Rambler on Android in select countries and languages, with Chrome next. Agora, Fishjam, LangChain, LiveKit, Pipecat, Vercel and Vision Agents support Gemini Live API development. Vivo, Intellitek Health and Lingopal cited latency, accuracy and language breadth positively.
Positives
- Artificial Analysis measured average word error rates of 4.0% for streaming and 2.6% for non-streaming transcription.
- Final transcription arrives 70% faster than with Google’s previous Chirp 3 model.
- More than 85 languages are automatically detected, including regional accents, dialects and live language changes.
- Sub-second, bidirectional Live API streaming supports interactive voice agents and real-time captioning.
- Rambler, Antigravity and Gemini for macOS combine transcription with voice editing, screen context and broader Gemini functions.
Risks & concerns
- Support beyond three speakers remains experimental, while standard speaker attribution applies only to prerecorded audio.
- Developer and enterprise editions remain in public preview, including access through the Gemini API and Gemini Enterprise Agent Platform.
- Rambler is restricted to select countries and languages, while Gemini for macOS consumer access is English-only.
- Chrome integration and Gemini Enterprise for Customer Experience are not yet available.
- Function calling for image generation and file analysis is currently limited to the Gemini app on macOS.