Gemini 3.8 Flash TTS

Generate lifelike AI speech with Gemini 3.8 Flash TTS — adjust delivery line by line, stage two-voice scenes, and reach 130 languages.

Gemini 3.8 Flash TTS
Shape every line's delivery on Flash TTS, or render high-volume audio affordably with Flash-Lite TTS
AI Video Prompt Generator

Support

Pro AI Tools

Explore elite tools including the Suno AI Music Generator.

placeholder hero

Gemini 3.8 Flash TTS Puts Directorial Control Into Text to Speech

Launched September 23, 2026, Google's newest speech release pairs an expressive creative model with a low-cost engine built for sheer volume.

  • Two Models, Two Very Different Jobs
    The flagship handles nuanced, character-driven reads; Flash-Lite keeps per-minute costs low when you need thousands of clips.
  • You Direct the Read Instead of Choosing a Preset
    Style prompts, structured speech metadata and inline vocal events work together to control mood, tempo, feeling and accent.
  • Design a Voice, or Clone One With Permission
    Sketch a persona in plain words, or mirror a real person using a reference clip plus their matching consent sample.

Getting Clean Audio Out of Gemini 3.8 Flash TTS

Four habits that keep your transcript readable and let the performance data do the acting.

Capabilities Behind Gemini 3.8 Flash TTS, Explained

Performance control, cloning safeguards and multilingual reach — a closer look at what this flagship voice model ships with.

Line-Level Performance Control

Per-turn style prompts plus inline laughs, sighs, coughs, breaths and beats feel less like picking a preset and more like coaching a performer.

Voices Designed in Plain English

Describe age, personality, accent, vocal texture and role, then draw on 2,000+ production voices served through the Voices endpoint.

Cloning That Requires Permission

A clean reference take plus a consent recording from the same adult speaker, with SynthID watermarking and C2PA credentials attached.

Two Speakers, One Continuous Take

Script podcast banter, classroom exchanges, product walkthroughs or game scenes without stitching separate lines together by hand.

Stable Tone Across Long Recordings

Google documents consistent identity, timbre, loudness and room tone across multi-minute narration and extended dialogue.

130 Languages, Accents Included

Flash TTS supports 130 languages against Flash-Lite's 101, with regional accents, minority dialects and IPA pronunciation overrides.

FAQ

Gemini 3.8 Flash TTS: Answers to What Users Ask Most

Quick replies on Gemini 3.8 Flash TTS cost per minute, tier choice, benchmark standings and voice-cloning rules.

1

What does Gemini 3.8 Flash TTS charge per minute of audio?

Roughly 1.35 cents per audio minute, based on launch pricing of $0.50 per million input tokens and $9 per million output tokens.

2

When should I choose Flash TTS over Flash-Lite TTS?

Reach for Flash TTS when acting nuance and long-form narration matter; Flash-Lite TTS wins on bulk jobs and latency-sensitive pipelines.

3

How does it stack up against rival voice models?

Google cites 71.4 on Hume's Voice Design Benchmark, and Voice Arena currently lists it in second place with 1,260 Elo.

4

What changed from the Gemini 3.1 Flash TTS Preview?

Flash-Lite TTS succeeds the 3.1 preview and drops audio output pricing from $20 to $6 per million tokens.

5

Can I clone a voice, and what safeguards apply?

Yes, with two recordings: a reference clip and a matching consent statement from the same adult speaker.

6

Why does the model read my stage directions out loud?

Because your input is treated as a verbatim transcript — move long-lasting directions into speech metadata instead.

Run Your Own Script Through Gemini 3.8 Flash TTS

Try both tiers inside Google AI Studio or the Gemini API — swapping between Gemini 3.8 Flash TTS and Flash-Lite TTS means editing one model identifier. Weigh batch against priority inference before locking in a production budget.