Gemini 3.8 Flash TTS
Generate lifelike AI speech with Gemini 3.8 Flash TTS — adjust delivery line by line, stage two-voice scenes, and reach 130 languages.
Support
Pro AI Tools
Explore elite tools including the Suno AI Music Generator.
MiniMax H3
MiniMax H3 AI Video Generator
Seedance 2.5
The Future of AI Video Is Here.

Seedance 2.0
The Future of AI Video Is Here.

Veo3.1
Create Stunning Videos with Veo3.1
FLUX 3 Video Generator

Kling 3.0
Next-Gen AI Video Generator
Grok Video Generator
Create Videos from Text or Images with AI
MiniMax H3 video generator
MiniMax H3 AI Video Generator

Gemini 3.8 Flash TTS Puts Directorial Control Into Text to Speech
Launched September 23, 2026, Google's newest speech release pairs an expressive creative model with a low-cost engine built for sheer volume.
- Two Models, Two Very Different JobsThe flagship handles nuanced, character-driven reads; Flash-Lite keeps per-minute costs low when you need thousands of clips.
- You Direct the Read Instead of Choosing a PresetStyle prompts, structured speech metadata and inline vocal events work together to control mood, tempo, feeling and accent.
- Design a Voice, or Clone One With PermissionSketch a persona in plain words, or mirror a real person using a reference clip plus their matching consent sample.
Getting Clean Audio Out of Gemini 3.8 Flash TTS
Four habits that keep your transcript readable and let the performance data do the acting.
Capabilities Behind Gemini 3.8 Flash TTS, Explained
Performance control, cloning safeguards and multilingual reach — a closer look at what this flagship voice model ships with.
Line-Level Performance Control
Per-turn style prompts plus inline laughs, sighs, coughs, breaths and beats feel less like picking a preset and more like coaching a performer.
Voices Designed in Plain English
Describe age, personality, accent, vocal texture and role, then draw on 2,000+ production voices served through the Voices endpoint.
Cloning That Requires Permission
A clean reference take plus a consent recording from the same adult speaker, with SynthID watermarking and C2PA credentials attached.
Two Speakers, One Continuous Take
Script podcast banter, classroom exchanges, product walkthroughs or game scenes without stitching separate lines together by hand.
Stable Tone Across Long Recordings
Google documents consistent identity, timbre, loudness and room tone across multi-minute narration and extended dialogue.
130 Languages, Accents Included
Flash TTS supports 130 languages against Flash-Lite's 101, with regional accents, minority dialects and IPA pronunciation overrides.
Gemini 3.8 Flash TTS: Answers to What Users Ask Most
Quick replies on Gemini 3.8 Flash TTS cost per minute, tier choice, benchmark standings and voice-cloning rules.
What does Gemini 3.8 Flash TTS charge per minute of audio?
Roughly 1.35 cents per audio minute, based on launch pricing of $0.50 per million input tokens and $9 per million output tokens.
When should I choose Flash TTS over Flash-Lite TTS?
Reach for Flash TTS when acting nuance and long-form narration matter; Flash-Lite TTS wins on bulk jobs and latency-sensitive pipelines.
How does it stack up against rival voice models?
Google cites 71.4 on Hume's Voice Design Benchmark, and Voice Arena currently lists it in second place with 1,260 Elo.
What changed from the Gemini 3.1 Flash TTS Preview?
Flash-Lite TTS succeeds the 3.1 preview and drops audio output pricing from $20 to $6 per million tokens.
Can I clone a voice, and what safeguards apply?
Yes, with two recordings: a reference clip and a matching consent statement from the same adult speaker.
Why does the model read my stage directions out loud?
Because your input is treated as a verbatim transcript — move long-lasting directions into speech metadata instead.
Run Your Own Script Through Gemini 3.8 Flash TTS
Try both tiers inside Google AI Studio or the Gemini API — swapping between Gemini 3.8 Flash TTS and Flash-Lite TTS means editing one model identifier. Weigh batch against priority inference before locking in a production budget.
