Google wants developers to stop picking voices off a shelf and start casting them. On Wednesday the company released Gemini 3.8 Flash TTS and a cheaper sibling, Gemini 3.8 Flash-Lite TTS. Together they offer a library of more than 2,000 production-ready voices, prompt-based voice design in more than 100 languages and dialects, and the ability to clone a voice from a 30-second sample once the voice's owner has recorded their consent.

In its announcement, Google said the two models are "transforming voice generation from static presets into a dynamic creative studio." The company's own numbers show how big the change is. Its earlier speech models shipped with 30 stock voices, and the blog post invites developers to "scale up from 30 original voices to an infinite library." Logan Kilpatrick, who leads Google's AI Studio and Gemini API developer push, called the release "our new SOTA text to speech model" on X. He listed voice design, voice replication, 100-language support and a first-place finish on Hume AI's voice benchmarks among its headline features.

Two models, two jobs

Google is splitting its speech line by workload. Flash TTS is the creative tier. The company pitches it for games, audiobooks, podcasts and interactive media. Users write a natural-language description of a voice's role, accent and character, and the model builds a matching voice. Demos include a Melbourne DJ and a fire-breathing Japanese dragon. The preset library covers regional varieties such as Mexican Spanish, Quebec French and Scots English. Custom voices can be saved and reused so they stay consistent across long projects. A "voice remixing" feature is announced but not yet available. It will let users adjust the timbre, pitch, pace and accent of library voices with prompts.

Flash-Lite TTS is the volume tier. It is built for dubbing, bulk audio content and real-time voice agents, where cost per hour counts for more than dramatic range. Both models share the new performance controls. Developers can write stage directions for every line, stage two-speaker scenes from a single script with natural turn-taking, and generate hours of continuous audio with what Google calls minimal speaker drift. Scripts can also include inline markup. Vocal bursts are written as tags such as `<laughs>`, `<sigh>` and `<gasp>`, and backchannel interjections as `|mhm|` or `|yeah|`, so writers can place reactions and comedic beats exactly.

Cloning, with guardrails

Voice replication is the feature most likely to draw scrutiny. Flash TTS can build a voice profile from 30 seconds of audio, but the user must first submit a spoken consent statement from the voice's owner. The system then checks that the consent recording matches the reference speaker before it creates the voice. Every clip from Google's Gemini Audio models carries an inaudible SynthID watermark, and replicated voices also carry C2PA content credentials. The feature is not available everywhere. A footnote in Google's post says voice replication through AI Studio is unavailable in Illinois, Texas, the European Economic Area, the UK, Switzerland and India. On quality, Google cites third-party leaderboards. Flash TTS took first place on Hume AI's Voice Design Benchmark with a score of 71.4 and also led the accent-modeling category at 60.8. Flash TTS and Flash-Lite TTS placed first and second on Hume's Overall Quality Index. Google says the models also reached top positions in blind preference tests on Voice Arena in Japanese, Brazilian Portuguese, Vietnamese, Modern Standard Arabic, Mexican Spanish and Hindi. Early hands-on testing was more mixed. The Decoder's Matthias Bastian prompted a heavily accented Berliner and got a convincing accent and intonation, but he also heard a high-pitched background whine in parts of both test clips.

Pricing is on the page, and it doubles in 2027

Google's API pricing page does list rates, as reported by The Decoder. Through the end of 2026, both models charge $0.50 per million text input tokens. Audio output costs $9 per million tokens on Flash TTS and $6 on Flash-Lite TTS. Google counts one second of audio as 25 tokens, so an hour of output comes to roughly $0.81 on Flash and $0.54 on Flash-Lite. On January 1, 2027, all of those rates double, to about $1.62 and $1.08 per hour. Both models are rolling out now in the Gemini API and Google AI Studio. Flash TTS is also coming to Gemini Notebook and Flash-Lite TTS to Google Vids, and Gemini Enterprise access is listed as coming soon. Launch partners include Figma, HeyGen and Wondercraft, and the models already work with Agora, LiveKit, Pipecat and Vercel.

Why It Matters

Voice has become one of the most contested layers of the generative AI stack, and the release puts pressure on companies whose main business is voice. ElevenLabs built its lead on voice cloning and a large voice marketplace. OpenAI sells steerable speech models through its API and uses voice heavily in ChatGPT. Alibaba's Qwen team has been releasing capable open-weight speech models that cost developers little or nothing. Google is now combining prompt-based voice design, consented cloning, multi-speaker long-form generation and an under-a-dollar-an-hour introductory price in one API. That combination turns features specialists once charged a premium for into something closer to a commodity.

The bigger strategic point is distribution. A voice model that is also inside Notebook, Vids and Gemini Enterprise, and already plugged into the main voice-agent frameworks, does not need to win every benchmark to win market share. Its consent-check and watermarking approach could also become a template rivals are pushed to match.

What to Watch

The January 2027 price doubling is the clearest test of whether developers stay once the introductory rates end, especially on the high-volume Flash-Lite tier where ElevenLabs and open-weight alternatives compete hardest on cost. Watch for the promised voice remixing feature, for enterprise availability through Gemini Enterprise, and for whether Google brings voice replication to the EU, UK and India, where it is currently blocked. Also watch whether independent testers confirm the benchmark rankings.

“transforming voice generation from static presets into a dynamic creative studio”
— Leland Rechis and Alan Cowen, Gemini Audio team, Google
2,000+
Production-ready voices in the library
71.4
Flash TTS score, #1 on Hume AI Voice Design
30 sec
Sample needed for consented voice cloning
$0.81/hr
Flash TTS audio output cost through 2026