Google released Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS on 23 September 2026, both rolling out the same day in the Gemini API and Google AI Studio. Google reports that Flash TTS takes the top overall position on Hume AI's Voice Design Benchmark at 71.4 and leads accent modeling at 60.8, with Flash-Lite TTS second on the same Overall Quality Index. Both cover 100+ languages and dialects and reach 2,000+ production-ready voices.

The number that decides whether you should build on it is not on the announcement page at all. It is on the Gemini API pricing page, and it comes with a date attached: audio output costs $9.00 per 1M tokens "through December 31, 2026" and $18.00 per 1M tokens "starting January 1, 2027". Flash-Lite moves from $6.00 to $12.00 on the same day. Every tier doubles.

That turns the headline into something narrower than it looks. Against the model it replaces, Gemini 3.1 Flash TTS Preview at $20.00 per 1M output tokens, Flash TTS is 55% cheaper today and 10% cheaper in January. The price cut is real, and most of it expires in 99 days.

What Google actually shipped

Two models, not one, and they are aimed at different jobs. Flash TTS is positioned for deep creative direction and character design. Flash-Lite TTS is positioned for high-volume, cost-efficient scale. Both landed in the Gemini API and Google AI Studio on launch day, with Gemini Enterprise listed as coming soon. Flash TTS also went into Gemini Notebook; Flash-Lite went into Google Vids.

The capability list is genuinely strong for anyone producing narration. You can create a custom voice from scratch or replicate one from a 30-second audio sample, and Google reports top positions in blind human preference evaluations on Voice Arena across Japanese, Brazilian Portuguese, Vietnamese, Modern Standard Arabic, Mexican Spanish and Hindi. Output is watermarked with SynthID.

One disambiguation worth making early, because Google shipped three different things under the same version number. Gemini 3.8 Flash TTS is text-to-speech. Gemini 3.8 Flash is the coding and agentic language model. Gemini 3.8 Live is the speech-to-speech model for real-time voice agents, which bills per minute rather than per token and sits in the same category as GPT-Live-1. They share a version string and nothing else, including their price lists.

Three charcoal plaques engraved Flash TTS, Flash and Live, three products sharing the Gemini 3.8 version number
Three different products ship under the Gemini 3.8 name, on three different price lists.

The price is a promotion, and the expiry is published

Google is unusually direct about this. The model card and the pricing table both carry the two-date structure, so nobody can claim they were not told. What the announcement does not do is put the January figure anywhere near the benchmark claim.

ModelAudio output, through 31 Dec 2026Audio output, from 1 Jan 2027Change
Gemini 3.8 Flash TTS$9.00 / 1M tokens$18.00 / 1M tokens+100%
Gemini 3.8 Flash-Lite TTS$6.00 / 1M tokens$12.00 / 1M tokens+100%
Text input, both models$0.50 / 1M tokens$1.00 / 1M tokens+100%
Gemini 3.8 Flash TTS, batch$4.50 / 1M tokens$9.00 / 1M tokens+100%
Gemini 3.1 Flash TTS Preview$20.00 / 1M tokens$20.00 / 1M tokensno change

Read the bottom row against the top one. The predecessor's price is stable; the successor's is not. A team that benchmarks Flash TTS in October against a budget built on $9.00, signs off, and ships in Q1 will be running on a number that no longer exists. The batch tier is the one genuine structural discount here, at exactly 50% off standard, and it survives the reset in relative terms even though it doubles in absolute ones.

Two charcoal platforms engraved $9.00 and $18.00 with a 1 Jan marker, the audio output price doubling
Audio output goes from $9.00 to $18.00 per 1M tokens on 1 January 2027, a 100% increase.

What a single maxed-out request costs

The model card gives the request envelope: text input up to 8K tokens, audio output up to 64K tokens. That is enough to price the worst case exactly, without estimating anything.

ConfigurationInput costOutput costTotal per request
Flash TTS, batch, now$0.0020$0.2880$0.2900
Flash-Lite TTS, standard, now$0.0040$0.3840$0.3880
Flash TTS, standard, now$0.0040$0.5760$0.5800
Flash-Lite TTS, standard, Jan 2027$0.0080$0.7680$0.7760
Flash TTS, standard, Jan 2027$0.0080$1.1520$1.1600
Gemini 3.1 Flash TTS Preview$0.0080$1.2800$1.2880

Two things fall out of that table. First, the input price is noise. At a full 8K tokens of text driving a full 64K tokens of audio, input is 0.69% of the bill and output is 99.31%. Any comparison that leads with the $0.50 input rate is measuring the wrong column.

Second, Flash-Lite is 33% cheaper than Flash, and that ratio holds after the reset because both double together. If Flash-Lite is good enough for your material, it is permanently a third cheaper, which is a more durable reason to pick it than the promotional window is.

Three vendors, three billing units, one uncomparable column

Here is where the real friction sits for anyone choosing a narration engine this week. Google bills audio by the token. OpenAI bills its classic TTS models by the character and its newer one by the token. ElevenLabs bills by the character, at one credit per character.

ModelBilling unitPublished priceNormalized where possible
Gemini 3.8 Flash TTSper 1M tokens, audio out$9.00 (then $18.00)not convertible
Gemini 3.8 Flash-Lite TTSper 1M tokens, audio out$6.00 (then $12.00)not convertible
OpenAI gpt-4o-mini-ttsper 1M tokens, audio out$12.00not convertible
OpenAI tts-1per 1M characters$15.00$15.00 / 1M chars
OpenAI tts-1-hdper 1M characters$30.00$30.00 / 1M chars
ElevenLabs v3 and v2 Multilingualper 1,000 characters$0.10$100.00 / 1M chars
ElevenLabs v3 Conversational, Flash, Turboper 1,000 characters$0.05$50.00 / 1M chars

The price column cannot be read downward. Within the per-character group the comparison is clean and the spread is large: ElevenLabs Multilingual at $100 per 1M characters is 6.7 times OpenAI tts-1 at $15, and ElevenLabs Flash at $50 is 3.3 times it. Across the groups, the comparison does not exist, because a token of audio and a character of text are not the same quantity and no published table converts between them.

Three mismatched charcoal measuring blocks engraved tokens, characters and credits, three incompatible billing units
Google bills audio by the token, OpenAI and ElevenLabs by the character. The column does not compare.

The number Google does not publish

Google documents exactly one audio token conversion. The token counting page states "Audio: 32 tokens per second", and it states it for audio input. There is no published tokens-per-second figure for audio output anywhere I could find across four Google surfaces: the announcement, the pricing table, the speech generation guide, and the model card. The speech generation guide specifies four encodings and three sample rates and no billing conversion. The model card explicitly carries no token counting methodology.

So the honest position is an inference, stated as one. If the output tokenizer runs at the same documented 32 tokens per second as input, then 1M output tokens is 31,250 seconds, or 8.68 hours of speech, and the arithmetic becomes:

  • Flash TTS today: about $1.04 per hour of generated audio
  • Flash TTS from January: about $2.07 per hour
  • Flash-Lite today: about $0.69 per hour
  • Flash-Lite from January: about $1.38 per hour
  • Flash TTS batch today: about $0.52 per hour

On the same assumption, the 64K output ceiling is about 33 minutes of audio in a single request. Those are good numbers, and they are cheap against the per-character vendors on any plausible speech rate. But they rest on a rate Google published for the other direction, which is why the first thing to do is measure it rather than trust this paragraph.

ElevenLabs does not close the gap from its side either. Its pricing page defines text-to-speech as one credit per character and publishes a minutes conversion for speech-to-text at 330 credits per minute, but no characters-to-minutes rate for generation. Neither vendor lets you answer "what does an hour of narration cost" from published material alone. This is the same pattern we found when a leaderboard position and the deployed serving reality diverged for Qwen3-TTS.

Charcoal plate engraved 32/sec beside a deliberately blank plate, the unpublished audio output token rate
Google documents 32 tokens per second for audio input. There is no published figure for output.

On the benchmark claim

The 71.4 and the 60.8 are Google's figures, attributed to Hume AI's benchmark. Hume's own landing page for Real World VoiceEQ describes six leaderboards covering speech recognition, speech understanding, text-to-speech quality, live voice agent performance, voice replication and voice controllability, but does not display the ranking table itself; the public leaderboard lives in a separate Hugging Face Space. That does not make the claim wrong. It does mean a reader cannot confirm the position from the benchmark's front page, and vendor-reported placements on third-party benchmarks deserve the same scepticism whoever reports them.

Which one to use

If you are producing long-form narration at volume, audiobooks, course modules, localized voiceover, start on Flash-Lite and batch it. That combination is the cheapest credible option on this list, it is a third below Flash permanently rather than temporarily, and Google's own positioning names high-volume scale as its job.

If you are doing character work, directed performance or accent-critical localization, Flash TTS is the model Google put the accent-modeling number on, and the 30-second voice replication path makes it the natural fit for consistent character voices across episodes.

Whichever you pick, do the measurement before the reset. Generate a representative sample of your actual script, read the billed output tokens off the response, divide by the audio duration you got back, and you will have the tokens-per-second rate for your content that no vendor publishes. Then price the January column, not the September one, because that is the number your production will actually run on.

Frequently asked questions

How much does Gemini 3.8 Flash TTS cost?

$0.50 per 1M text input tokens and $9.00 per 1M audio output tokens through 31 December 2026, rising to $1.00 and $18.00 on 1 January 2027. Flash-Lite TTS is $6.00 per 1M audio output tokens now and $12.00 from January. Batch pricing is exactly half the standard rate in each case.

Why does the price double on 1 January 2027?

Google published the launch rates as time-limited. The pricing table carries both figures with the wording "through December 31, 2026" and "starting January 1, 2027" across every service tier, including Standard, Batch, Flex and Priority. It is a promotional launch price with a stated expiry, not a price cut.

Is Gemini 3.8 Flash TTS cheaper than ElevenLabs or OpenAI?

You cannot answer that from the published prices, because Google bills audio output per token while ElevenLabs and OpenAI's tts-1 models bill per character, and no vendor publishes a conversion between the two. Within the per-character group the comparison works: OpenAI tts-1 is $15 per 1M characters against ElevenLabs Multilingual at $100 per 1M characters. To compare across groups you have to measure your own token-to-duration ratio.

What is the difference between Gemini 3.8 Flash TTS and Gemini 3.8 Live?

Flash TTS is text-to-speech: you send text, you get audio. Gemini 3.8 Live is speech-to-speech for real-time voice agents, it bills per minute rather than per token, and it sits on a different price list. They share a version number and little else. Gemini 3.8 Flash, without a suffix, is a third thing again: the coding and agentic language model.

How long can a single Gemini 3.8 TTS request be?

The model card gives text input up to 8K tokens and audio output up to 64K tokens. If output tokens are counted at the 32 tokens per second that Google documents for audio input, 64K tokens is roughly 33 minutes of audio per request, but Google does not publish an output rate, so treat that as an inference to verify against your own content.

Should I switch from Gemini 3.1 Flash TTS Preview?

On price, yes, and the case is strongest before January. The 3.1 preview is $20.00 per 1M output tokens with no scheduled change, so Flash TTS at $9.00 is 55% cheaper today and still 10% cheaper after the reset, while Flash-Lite at $6.00 stays 40% cheaper even at its January rate of $12.00.

Does Gemini 3.8 TTS watermark its output?

Yes. Google states that output carries SynthID watermarking. If your distribution pipeline re-encodes or heavily processes the audio, test whether the watermark survives your chain before assuming provenance signalling is intact downstream.