Google officially launched Gemini 3.8 Flash TTS and Flash-Lite TTS on September 23, 2026.
Google opens Gemini speech models to direct voice design
Software developers gain natural language tools to build custom artificial voices instead of picking from preset lists.
In a nutshell
Google has released two text-to-speech engines that let software developers design realistic voices using plain written instructions rather than static catalogues. The system handles emotional acting cues across dozens of languages, immediately drawing commercial software makers to automate voice production while embedding mandatory digital watermarks to identify computer-generated sound.
Highlights
- Google released two dedicated speech systems spanning Flash TTS and the lower-cost Flash-Lite TTS.
- The flagship Flash TTS system supports script directing and conversational reactions across 130 languages.
- The streamlined Flash-Lite TTS model provides high-volume vocal synthesis across 101 languages.
- Gemini 3.8 Flash TTS topped Hume AI's Voice Design Benchmark with an evaluation score of 71.4.
- The platforms include a catalog of over 2,000 preset voices with mandatory SynthID watermarking on all outputs.
Gemini voice model language coverage
| Model tier | Languages supported |
|---|---|
| Flash TTS | 130 |
| Flash-Lite TTS | 101 |
- The total number of languages supported across Google's two new synthetic speech models.
- The count determines which international markets and dubbing pipelines can immediately deploy the technology.
From the Editor’s Diary
When synthetic voices can be directed like human actors through plain written text, voice generation changes from a simple playback feature into an active design layer across media software.
Who's involved
Google
Multinational technology company developing artificial intelligence software
goal → Drive business adoption of its audio tools against competing synthetic voice services
Leland Rechis
Group Product Manager supervising speech tools at Google
goal → Direct the release of programmable voice technology to software engineers and media companies
Hume AI
Artificial intelligence research firm tracking emotional realism in synthetic speech
goal → Score the expressive performance and authenticity of competing corporate voice tools
Figma and HeyGen
Visual design and video generation companies adopting Google software
goal → Integrate custom voice prompts and automated language dubbing into production software
In short
Google has turned computer voice generation into a steerable design tool, letting software developers describe how a speaker should sound rather than choosing from a fixed menu of canned voices. The release removes an old bottleneck in automated media production by allowing engineers to control pacing, acting directions, and accents through regular typed text.
The rollout will likely accelerate the replacement of human voice actors and customer service operators with programmable synthetic voices across games, video dubbing, and corporate call centers.
That shift is likely to happen quickly because established design software makers and commercial dubbing applications have already plugged the technology directly into active customer workflows.
How it unfolded
Google announces Gemini 3.8 Flash TTS
Google product managers led by Leland Rechis launched the Gemini 3.8 Flash TTS and Flash-Lite TTS models on September 23, 2026, pitching them as steerable software engines rather than simple digital playback tools. The announcement gave outside programmers an interactive testing ground and direct code hooks to shape speech delivery with everyday written prompts.
API rollout and benchmark evaluations
The technological shift became clear when Google released technical manuals separating the plain text of a script from vocal cues like laughs, sighs, and emotional direction tags. The premium model promptly scored 71.4 points to lead Hume AI's Voice Design Benchmark for expressive realism, prompting immediate integration announcements from commercial video and design services including Figma, HeyGen, and Wondercraft.
Where things stand
The two speech engines operate live in Google AI Studio and developer coding interfaces, while built-in features are rolling out across Google Vids and Gemini Notebook. Expanded corporate access and an interactive voice-remixing tool are set to follow in later updates.