Google has introduced two new text-to-speech models, Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, through its cloud platform. These models deliver top-tier audio quality and speed improvements while supporting extensive language and voice customization capabilities.

  • Two new text-to-speech models launched with similar APIs for ease of integration
  • Flash TTS supports 130 languages with premium audio quality; Flash-Lite TTS targets cost and speed with 101 languages
  • Models include advanced voice customization and AI-generated voice watermarking for security

What happened

Google rolled out two new text-to-speech models named Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS via its cloud platform. Both models offer nearly identical APIs, facilitating straightforward developer adoption and side-by-side usage. Flash-Lite TTS aims for cost efficiency and faster processing, while Flash TTS delivers superior audio fidelity at a premium pricing tier.

At launch, Flash TTS can generate speech in 130 languages, while Flash-Lite supports 101. Users can access a vast library of more than 2,000 prepackaged voices and can customize voice parameters such as timbre, accent, and pacing. Developers are also able to create synthetic voices from short audio samples with explicit speaker consent, backed by embedded watermarking technology to ensure traceability.

Why it matters

These models enable a wide range of applications requiring high-quality and flexible speech generation, such as audiobooks, virtual assistants, and video content narration. The improved latency and cost benefits offered by Flash-Lite make advanced speech tech accessible to a broader developer community and enterprise use cases that require efficiency at scale.

Google’s inclusion of audio watermarking through SynthID and C2PA digital provenance records reflects growing concerns about AI-generated content authenticity and security. This positions Google’s models not only as powerful tools for creators but also as responsible implementations in the evolving AI landscape. Their benchmark-topping results further underline the company’s competitive lead in speech synthesis technology.

What to watch next

Google plans to expand voice customization options beyond current capabilities, including the ability to modify existing prepackaged voices to create new ones. Monitoring how developers adopt these models in real-world applications and the impact on cloud speech usage metrics will be key indicators of success.

The broader rollout of SynthID watermarking and C2PA records across audio-generated content will also be important to observe, as they may set standards for AI voice content provenance and detection. Additionally, competition from other cloud providers and startups in text-to-speech technology will likely prompt further innovation and benchmarking shifts in the near term.

Source assisted: This briefing began from a discovered source item from SiliconANGLE. Open the original source.
How SignalDesk reports: feeds and outside sources are used for discovery. Public briefings are edited to add context, buyer relevance and attribution before they are published. Read the standards

Related briefings