Gemini 3.8 Flash TTS: Google’s New Voice AI Can Design, Clone, and Direct Voices — Here’s How It Works
On September 23, Google released Gemini 3.8 Flash TTS and Flash-Lite TTS — text-to-speech models that let you describe a voice in words and get it, replicate a voice from a 30-second sample, and direct delivery line by line like a voice actor. Here’s what they do and whether they’re worth your attention.
What’s in this guide
- Two models, two jobs
- The key features
- Safety: consent and watermarking
- Where you can use it
- Who it’s for
- FAQ
Two Models, Two Jobs
| Model | Built for | Consumer access |
|---|---|---|
| Gemini 3.8 Flash TTS | Creative direction, characters, designing new voices from prompts | Gemini Notebook |
| Gemini 3.8 Flash-Lite TTS | High-volume, cost-efficient output: dubbing, voice agents | Google Vids |
Google calls these its “most expressive audio models yet.” Both are available to developers now via the Gemini API and Google AI Studio; enterprise access through Gemini Enterprise is coming soon.
The Key Features
Voice design
Describe a voice in natural language — age, tone, accent, energy — and create it from scratch.
2,000+ voices
A library of production-ready voices across 100+ languages and dialects.
Voice replication
Recreate a consistent voice from a 30-second sample, with verbal consent from the voice owner.
Line-by-line direction
Control pacing, emotion, and delivery with stage directions or script cues.
Two-speaker scenes
Generate natural back-and-forth dialogue, useful for podcasts and training content.
Vocal textures
Add laughs, sighs, and listening sounds like “mm-hmm” for more human-sounding audio.
Google says the models hold quality across hours of continuous audio, which matters for audiobooks and long-form narration, and that they took the top spot on Hume AI’s voice design benchmark. Google’s Logan Kilpatrick also teased “voice remixing” as coming soon.
Safety: Consent and Watermarking
Voice cloning is one of the riskiest areas of generative AI — voice-based fraud is a growing concern. Google has added two guardrails. First, voice replication requires verbal consent from the person whose voice is being copied. Second, all generated audio carries SynthID, an imperceptible watermark built into the audio that allows it to be identified as AI-generated.
Neither is a complete fix: watermarks help platforms detect AI audio, but a scam call won’t be checked by the person receiving it. Businesses and families should still use verification steps, like call-back rules or code words, for any urgent request involving money.
Where You Can Use It
- Developers: Gemini API and Google AI Studio, available now.
- Everyday users: Flash TTS in Gemini Notebook, and Flash-Lite TTS in Google Vids for video voiceovers.
- Enterprises: Gemini Enterprise, coming soon.
Who It’s For
- Content creators: voiceovers for YouTube, courses, and social video without recording yourself.
- Businesses with voice agents: more natural-sounding phone and support bots, at scale with Flash-Lite.
- Localization teams: dubbing into 100+ languages with a consistent voice.
- Podcasters and educators: two-speaker dialogue for explainers and lessons.
FAQ
Can I clone anyone’s voice?
No. Voice replication requires verbal consent from the voice owner, and all output is watermarked.
How many languages does it support?
More than 100 languages and dialects, according to Google.
Do I need to code to use it?
No. You can try it in Google AI Studio, or use it inside Gemini Notebook and Google Vids.
Related Reading on FutureLume
- AI Voice Cloning in 2026: Legitimate Uses vs. the Fraud Risk
- Best AI Video Generators in 2026: Veo vs Kling vs Seedance vs Sora
- GPT-6 Astra vs Claude Fable 5.1 vs Gemini 3.8 Flash vs Muse Spark 1.3: September 2026’s AI Model Launches Compared
Gemini 3.8 Flash TTS moves AI voice from “pick a voice from a list” to “direct a performance.” For creators and businesses already in Google’s ecosystem, it’s an easy upgrade to try. Just use voice cloning responsibly — the consent requirement exists for good reason.
