Google Unveils Gemini 3.8 Flash TTS – AI Creates Voices From a Description and Reproduces Them From a Short Sample

Google has released two new speech synthesis models – Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS. The company calls them its most expressive audio generation models yet: the flagship version is built for complex voiceover work, acting and character creation, while Flash-Lite targets mass dubbing, text reading and voice agents.
One of the key features of Gemini 3.8 Flash TTS is Generative Voice Design. A developer can describe the voice they need in plain text – its age, timbre, manner of speech, accent and performance style – and the system will create a new voice persona.
Google has also opened up a library of more than 2,000 ready-made voices, including regional variants such as Mexican Spanish, Quebec French and Scottish English.
The flagship model supports 130 languages, while Flash-Lite supports 101. Gemini 3.8 Flash TTS can also handle regional accents and dialects and supports IPA phonetic hints for difficult pronunciations. The model’s specs are published in the official Gemini API documentation.
A voice can be reproduced from a short recording, but only with the owner’s consent
Both models support Voice Replication – creating a stable synthetic voice from a short audio sample.
This requires 10–30 seconds of clean speech, plus a separate recording of the same person including a mandatory phrase consenting to the creation of a synthetic model of their voice. The system checks that the two recordings match before creating the profile.
So you can’t just upload someone else’s video or an interview clip and clone their voice without the owner’s involvement using Google’s standard tools.
Once verified, the voice can be saved under a permanent voice_id and reused in projects. For scenarios where a developer doesn’t want to store the voice profile on Google’s servers, there’s also an option with a temporary encrypted key. Detailed requirements are described on the Voice Replication page.
introducing Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, our most expressive audio generation models yet
— Google AI Studio (@GoogleAIStudio) September 23, 2026
these models enable creators, developers, and enterprises to create richer, more expressive audio experiences
try them via the Gemini API and in AI Studio:… pic.twitter.com/vV84ErmKY1
Voiceover can be directed line by line
Gemini 3.8 Flash TTS lets you set individual instructions for virtually every line – pace, emotional coloring, intonation, accent and acting delivery.
The model also understands embedded events such as laughter, a sigh or a short pause, as well as natural conversational reactions. It supports two-character scenes, where the system has to maintain distinct voices and a natural back-and-forth of lines.
Google also claims improved stability in long audio recordings – voice, timbre, volume and acoustic environment should drift less over a lengthy dialogue or narration. This matters especially for audiobooks and podcasts.
A Voice Remixing feature is coming in the future – the ability to take an existing voice from the library and separately change its timbre, pitch, speed or accent using text commands. For now, the feature has only been announced.
Google claims leadership in certain voice benchmarks
In its official announcement, Google calls Gemini 3.8 Flash TTS its most expressive TTS model and cites the results of several tests.
In the Hume AI Voice Design Benchmark, the model took first place with 71.4 points and also posted the best result in accent evaluation – 60.8 points. Flash TTS and Flash-Lite took first and second place, respectively, in the Hume AI Overall Quality Index.
These results apply to specific tests and do not mean the model is absolutely superior to all competitors in every scenario.
To guard against abuse, every piece of audio created by Gemini Audio gets an invisible SynthID watermark. Voice replication additionally uses consent verification and C2PA metadata.
Gemini 3.8 Flash TTS and Flash-Lite TTS are already available to developers through the Gemini API and Google AI Studio. The flagship model is also coming to Gemini Notebook, and Flash-Lite to Google Vids. Voice replication through AI Studio is not yet available in a number of regions, including the EEA, the UK, Switzerland and India.