Back to wire
AI·Article

Google launches Gemini 3.8 Flash TTS with prompt-designed and replicated voices

Google has begun rolling out Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS through the Gemini API and Google AI Studio, adding prompt-designed voices, 30-second voice replication with matched verbal consent and line-by-line performance control. Paid standard API pricing through 31 December is $0.50 per million text input tokens and $9 per million audio output tokens for Flash, or $6 for Flash-Lite; Gemini Enterprise API access is still coming soon.

Published 23 Sept 2026, 21:40

Two TTS tiers move into the Gemini API

Google launched Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS on 23 September and began rolling both models through the Gemini API and Google AI Studio. Flash is the higher-fidelity tier for long-form narration, multi-speaker scenes, regional accents and detailed acting direction; Flash-Lite is tuned for lower-cost, high-throughput work such as dubbing, read-aloud features and voice-agent cascades.

The release expands Google’s speech stack beyond fixed voice presets. Both models accept structured per-turn direction for pacing and delivery, while Flash adds a broader creative voice-design layer. Google documents 130 supported languages for Flash and 101 for Flash-Lite; its launch post also advertises a library of more than 2,000 production-ready voices.

Voice design and replication make consent part of the workflow

Gemini 3.8 Flash TTS can create a voice from natural-language instructions describing characteristics such as role, accent and delivery. For voice replication, Google says a user can supply a roughly 30-second reference recording of their own voice or a voice they have rights to use, but the system also requires a verbal consent recording from the voice owner that matches the reference speaker before the replicated voice is created.

Google says every clip generated by Gemini Audio carries its SynthID watermark, while replicated voices are also backed by C2PA credentials. Those mechanisms provide provenance signals around generated speech, but the launch material does not establish how reliably consent checks resist determined impersonation attempts or how robust watermark detection remains after common audio transformations.

Pricing favours Flash-Lite for volume, with Enterprise access later

For standard paid Gemini API use through 31 December 2026, Google lists both TTS models at $0.50 per million text input tokens. Audio output costs $9 per million tokens for Flash and $6 for Flash-Lite; Google says audio is counted at 25 tokens per second. From 1 January 2027 those standard rates are scheduled to rise to $1 for input and $18 or $12 respectively for audio output. Batch and Flex tiers currently halve the standard input and output rates.

The rollout remains split by product. Developers can use both models in the Gemini API and AI Studio now, Flash also appears in Gemini Notebook, and Flash-Lite is available through Google Vids. Gemini Enterprise API access is listed as coming soon. Google separately says voice replication through AI Studio is unavailable in Illinois, Texas, the European Economic Area, the United Kingdom, Switzerland and India, so that particular interface should not be treated as globally available.

Launch benchmarks and safety claims still need outside testing

Google cites Hume AI and Voice Arena results in which the new models rank highly for voice design and perceived speech quality, but those launch-selected measurements do not establish performance across every language, speaker style or production pipeline. The DeepMind model card also lists familiar foundation-model limitations, including possible hallucinations, occasional slowness or timeouts and continuing work on jailbreak resistance.

The product release is clearest on what has shipped: two API-accessible speech models, a richer voice-design system, consent-gated replication and provenance tooling. How well those safeguards hold up against abuse, how consistently quality carries across the advertised language range, and whether the reported benchmark advantages survive independent reproduction remain open questions beyond Google’s release evidence.

Source trail

3 sources · 3 primary