Gemini's New TTS Models Treat Voiceover Like a Screenplay

Google's new Gemini TTS models let you design a voice from a description and direct dialogue line by line, in a screenplay-style editor.

Share
Gemini's New TTS Models Treat Voiceover Like a Screenplay

Google released Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS today, two text-to-speech models it calls its most expressive yet.

The feature that stands out for filmmakers isn't the voice quality alone. It's how you direct them: line by line, in a screenplay-style editor, with performance cues written straight into the script.

Two Models, Two Different Jobs

Flash TTS is built for creative production, including character voices, audiobooks, podcasts, and games. It's the model for designing a specific vocal persona and controlling its delivery in detail.

Flash-Lite TTS is the cheaper, high-volume option. Google positions it for dubbing, audio content at scale, and voice agents that adjust tone and pacing on the fly.

Both models are rolling out now in the Gemini API and Google AI Studio. Flash TTS is also coming to Gemini Notebook, and Flash-Lite TTS to Google Vids, with Gemini Enterprise access listed as coming soon.

Directing a Voice Like a Performance

Google's AI Studio now includes an audio playground built like a voice design workspace. You can describe a new voice in plain language, its role, accent, and character, or pick from more than 2,000 ready-made voices across 100+ languages.

From there, a dual-speaker screenplay editor lets you stage back-and-forth dialogue and adjust delivery one line at a time. Non-verbal cues like <laughs> or an active-listening "mhm" go directly into the text, and Google says the models hold consistent across hours of generated audio.

That's a meaningfully different workflow from typing a paragraph and hoping the read lands. It's closer to giving an actor notes than configuring a preset.

Voice Cloning, With Limits

Flash TTS can also recreate a voice from a 30-second sample, after a consent check. Every clip generated by Gemini's audio models carries a SynthID watermark, and cloned voices also get C2PA content credentials identifying them as AI-generated.

One real limitation: voice replication through AI Studio isn't available in Illinois, Texas, the European Economic Area, the UK, Switzerland, or India. If you're working in any of those places, voice design from a description still works, but cloning a real voice through AI Studio doesn't.

Where Filmmakers Could Actually Use This

The practical use cases map closely onto real production needs. Temp voiceover for an edit before booking talent, scratch dialogue for animatics and previs, narration drafts for documentary cuts, and localized dubs for festival or international distribution are all realistic starting points.

None of that replaces a real performer on a final mix. But for the stages of a project where you need a voice now and a booked actor later, direction-level control is the difference between a placeholder you can work with and one you fight against.

Competitive Context

This puts Google in direct competition with ElevenLabs, the established leader in AI voice, along with xAI's recently launched Grok Voice. Google says Flash TTS placed first on Hume AI's Voice Design Benchmark, and that both models ranked first and second on Hume AI's Overall Quality Index. Those are Google's own reported results, worth testing against your own material before switching tools.

Pricing is the other pressure point. Reported rates put Flash TTS audio output at $9 per million tokens through the end of 2026, down from $20 on the preview model it replaces.

The Signal in the Noise

AI voice tools have mostly competed on how natural a voice sounds. Google is betting that the next differentiator is how well you can direct it, treating voice generation as a performance you shape rather than a file you export.

For filmmakers, that's the right question to be asking of any voice tool. A realistic voice you can't direct is still a stock asset.

Would line-by-line direction change how you'd use AI voice in your own work, or does temp VO still feel like something you'd rather record yourself?

Specs & Pricing

  • Models: Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, released September 23, 2026
  • Flash TTS: creative production, character voices, audiobooks, podcasts, games
  • Flash-Lite TTS: lower-cost, high-volume dubbing, content, and voice agents
  • Voices: design from a text description, or choose from 2,000+ ready-made voices across 100+ languages
  • Direction: dual-speaker screenplay editor, line-by-line delivery control, inline cues like <laughs>
  • Voice cloning: from a 30-second sample with a consent check; not available via AI Studio in Illinois, Texas, the EEA, UK, Switzerland, or India
  • Safeguards: SynthID watermark on all output; C2PA credentials on cloned voices
  • Pricing: Flash TTS audio output reported at $9 per million tokens through 2026; Flash-Lite TTS priced lower
  • Access: Gemini API and Google AI Studio now; Gemini Notebook (Flash TTS) and Google Vids (Flash-Lite TTS); Gemini Enterprise coming soon

Resources & Reads