Google Gemini 3.8 TTS Launches Custom Voices in 100-Plus Languages
Google launches Gemini 3.8 Flash and Flash-Lite TTS with custom voices, 100-plus languages, consent checks, SynthID and broad tools for creators and enterprises.
Summary
On September 23, 2026, Google launched Gemini 3.8 Flash TTS for detailed creative direction and Gemini 3.8 Flash-Lite TTS for cost-efficient, high-volume dubbing, content and voice agents. Flash expands 30 original voices into unlimited custom designs, supports natural-language control of role, accent and vocal traits across more than 100 languages and dialects, and offers over 2,000 production-ready voices. Both models provide line-level control of acting, pacing, dialect, tone, two-speaker scenes and conversational sounds, plus hours-long generation with minimal drift. Users can save voices, while 30-second samples can replicate voices they own or may use. Voice remixing is coming soon.
Replication requires a matching verbal consent recording, with C2PA credentials, while every Gemini Audio clip receives an imperceptible SynthID watermark. Gemini 3.8 Flash scored 71.4 overall and 60.8 for accents to rank first on Hume AI’s Voice Design Benchmark; Flash and Flash-Lite placed first and second on its Overall Quality Index and led Voice Arena preferences in Japanese, Brazilian Portuguese, Vietnamese, Modern Standard Arabic, Mexican Spanish and Hindi. Both are available through the Gemini API and Google AI Studio, with Flash in Gemini Notebook and Flash-Lite in Google Vids. Gemini Enterprise API access is coming soon. Replication in AI Studio excludes Illinois, Texas, the EEA, UK, Switzerland and India. Agora, LiveKit, Pipecat and Vercel support deployment, while Figma, HeyGen, Linguana, Wondercraft, 99.co and Ollang are integrating the models.
Positives
- More than 100 languages and dialects can be shaped by role, accent and vocal characteristics through natural-language prompts.
- Over 2,000 production-ready voices supplement an effectively unlimited library of custom voice designs.
- A 71.4 overall score and 60.8 accent score put Gemini 3.8 Flash first on Hume AI’s Voice Design Benchmark.
- Thirty-second voice replication requires matching verbal consent and adds SynthID watermarking and C2PA credentials.
- Hours-long generation, two-speaker staging and line-level direction support audiobooks, podcasts, dubbing and interactive voice agents.
Risks & concerns
- Voice replication can reproduce a vocal profile from only 30 seconds of audio, making effective consent enforcement essential.
- AI Studio voice replication is unavailable in Illinois, Texas, the EEA, UK, Switzerland and India.
- Gemini Enterprise API access and voice remixing were announced as coming soon rather than available at launch.
- Long-form generation promises minimal speaker drift, not its complete elimination across hours of audio.