Confirmed

Google officially launched Gemini 3.8 Flash TTS and Flash-Lite TTS on September 23, 2026.

AI Models

Google opens Gemini speech models to direct voice design

Software developers gain natural language tools to build custom artificial voices instead of picking from preset lists.

Published
NRB — News Republic Brigade

Other links

Google

In a nutshell

Google has released two text-to-speech engines that let software developers design realistic voices using plain written instructions rather than static catalogues. The system handles emotional acting cues across dozens of languages, immediately drawing commercial software makers to automate voice production while embedding mandatory digital watermarks to identify computer-generated sound.

Highlights

  • Google released two dedicated speech systems spanning Flash TTS and the lower-cost Flash-Lite TTS.
  • The flagship Flash TTS system supports script directing and conversational reactions across 130 languages.
  • The streamlined Flash-Lite TTS model provides high-volume vocal synthesis across 101 languages.
  • Gemini 3.8 Flash TTS topped Hume AI's Voice Design Benchmark with an evaluation score of 71.4.
  • The platforms include a catalog of over 2,000 preset voices with mandatory SynthID watermarking on all outputs.

Gemini voice model language coverage

Model tierLanguages supported
Flash TTS130
Flash-Lite TTS101
  • The total number of languages supported across Google's two new synthetic speech models.
  • The count determines which international markets and dubbing pipelines can immediately deploy the technology.

From the Editor’s Diary

When synthetic voices can be directed like human actors through plain written text, voice generation changes from a simple playback feature into an active design layer across media software.

Who's involved

  • Google

    Multinational technology company developing artificial intelligence software

    goal → Drive business adoption of its audio tools against competing synthetic voice services

  • Leland Rechis

    Group Product Manager supervising speech tools at Google

    goal → Direct the release of programmable voice technology to software engineers and media companies

  • Hume AI

    Artificial intelligence research firm tracking emotional realism in synthetic speech

    goal → Score the expressive performance and authenticity of competing corporate voice tools

  • Figma and HeyGen

    Visual design and video generation companies adopting Google software

    goal → Integrate custom voice prompts and automated language dubbing into production software

In short

Google has turned computer voice generation into a steerable design tool, letting software developers describe how a speaker should sound rather than choosing from a fixed menu of canned voices. The release removes an old bottleneck in automated media production by allowing engineers to control pacing, acting directions, and accents through regular typed text.

The rollout will likely accelerate the replacement of human voice actors and customer service operators with programmable synthetic voices across games, video dubbing, and corporate call centers.

That shift is likely to happen quickly because established design software makers and commercial dubbing applications have already plugged the technology directly into active customer workflows.

How it unfolded

01

Google announces Gemini 3.8 Flash TTS

2026-09-23 – 2026-09-23

Google product managers led by Leland Rechis launched the Gemini 3.8 Flash TTS and Flash-Lite TTS models on September 23, 2026, pitching them as steerable software engines rather than simple digital playback tools. The announcement gave outside programmers an interactive testing ground and direct code hooks to shape speech delivery with everyday written prompts.

1 source
02

API rollout and benchmark evaluations

2026-09-23 – 2026-09-24

The technological shift became clear when Google released technical manuals separating the plain text of a script from vocal cues like laughs, sighs, and emotional direction tags. The premium model promptly scored 71.4 points to lead Hume AI's Voice Design Benchmark for expressive realism, prompting immediate integration announcements from commercial video and design services including Figma, HeyGen, and Wondercraft.

2 sources

Where things stand

The two speech engines operate live in Google AI Studio and developer coding interfaces, while built-in features are rolling out across Google Vids and Gemini Notebook. Expanded corporate access and an interactive voice-remixing tool are set to follow in later updates.

Sources