Skip to main content

Gemini 3.8 TTS Goes GA With 150 Voices and New Pricing

4 min read

Gemini 3.8 TTS Flash and Flash-Lite are generally available with voice design, consent-verified replication and more than 150 voices. Compare costs and uses.

Gemini 3.8 TTS Goes GA With 150 Voices and New Pricing

Google prices its new speech models by generated audio tokens. That makes your voice choice a creative decision and the length of every generated clip a budget decision.

Gemini 3.8 Flash TTS and Flash-Lite TTS became generally available through the Gemini API on September 22, 2026. The release also adds a Voices endpoint. Google documents voice design, consent-verified voice replication and a library of more than 150 voices. These are text-to-speech models: they take text and return generated audio. They are separate from the Gemini Live API, which is built for interactive audio conversations.

Two models for different production jobs

The model IDs are gemini-3.8-flash-tts : and gemini-3.8-flash-lite-tts. Google positions Flash for expressive, long-form voice work and Flash-Lite for faster, higher-volume generation. Descriptions such as “studio-grade” are Google’s product claims; teams should compare output with their own scripts, languages, voices, and listening conditions.

ChoiceBest first testSeptember 2026 paid standard audio output
Gemini 3.8 Flash TTSLong narration, character direction and expressive speech$9 per million audio tokens
Gemini 3.8 Flash-Lite TTSHigh-volume prompts, short answers and latency-sensitive cascades$6 per million audio tokens
Google’s published standard rates through December 31, 2026. Text input and other applicable charges are separate.

What a minute of generated speech can cost

Google’s pricing page says generated audio corresponds to 25 audio tokens per second. One minute therefore represents about 1,500 audio tokens. At the published standard paid output rates, that is about $0.0135 for Flash or $0.009 for Flash-Lite per generated minute, before text-input charges, retries, storage, or other usage. Ten thousand one-minute clips would put audio output alone at roughly $135 or $90 respectively. This is a MustHave.ai calculation from Google’s token conversion and rate card, not a quoted all-in price.

The introductory paid rates end December 31, 2026. Google’s listed standard audio-output rates rise to $18 and $12 per million tokens on January 1, 2027. A production forecast should use the rates for the month in which requests will run, not only the launch price.

Voice design and replication require separate review.

Voice design creates a persistent custom vocal persona from a description. Voice replication uses a person’s voice and, according to Google, requires consent verification. The technical ability to produce a similar voice does not grant permission to use someone’s likeness in a public product. Keep records of who authorized the source voice, for what purpose, and for how long. Review the actual license and policy terms before commercial deployment.

The Voices endpoint can help a team enumerate available voices rather than hard-code a name found in a demo. For a multilingual service, test pronunciation, names, interruptions, and long passages in every target language. A voice that sounds strong in one short sample may not hold up across a full lesson or support flow.

TTS is not the Live API

A voice agent can use TTS as the last step in a cascade: speech recognition, a text model, then generated audio. Gemini Live handles real-time audio interaction through a different model and interface. If your product needs quick back-and-forth conversation, measure end-to-end turn latency, not only TTS generation speed. If it produces a podcast, tutorial, or narrated video, listen for consistency across longer segments and edits.

A migration check for the 3.1 preview.

  1. Pin the new model ID in a test environment rather than changing a global alias.
  2. Run the same scripts through the preview and both 3.8 models.
  3. Compare speaking rate, pauses, pronunciation, voice stability, and output length.
  4. Count generated audio tokens and failed or repeated attempts.
  5. Test the voice-consent process before enabling replication for users.
  6. Roll out by language and use case, with a way to revert if quality drops.

For a related creator workflow, see our YouTube Gemini editor and voice-likeness report. Our speech-transcription pricing guide covers the other end of a voice pipeline: turning audio back into text.

Primary sources

Checked September 24, 2026. Google reports the model capabilities; the per-minute figures are calculated from its published audio-token conversion and paid standard rates.

Leave a comment

Your email address will not be published. Required fields are marked *