Gemini 3.8 Flash TTS, Gemini 3.8 Flash-Lite TTS Introduced by Google With Custom Voice Design and SynthID Watermarking
Google has launched Gemini 3.8 Flash TTS and Flash-Lite TTS models, introducing custom voice design, line-by-line performance direction, and voice replication from 30-second samples. Backed by SynthID watermarking and strict consent verification, the models are now available in Google AI Studio and the Gemini API.
Google has expanded its artificial intelligence lineup by introducing two new text-to-speech models to the Gemini family, transforming voice generation into a dynamic creative studio. The new offerings, Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, are designed to give developers, creators, and enterprises advanced tools to produce rich, expressive audio experiences across various digital platforms.
The newly released systems aim to elevate standard voice generation by moving beyond static presets into customizable audio production. As per a report by Google, the models integrate deeply into the Gemini ecosystem, improving applications such as Gemini Notebook and Google Vids while supporting over 100 languages and regional dialects worldwide. Adobe Premiere Video Editor Now Available on Google Play; Know What It Offers.
Custom Voice Design and Replication
The Gemini 3.8 Flash TTS model functions as a comprehensive vocal studio, allowing users to create entirely original character voices from scratch using natural language prompts. Creators can adjust roles, accents, and emotional tones to bring fictional characters to life for audiobooks, podcasts, and interactive media. Furthermore, the system includes a voice replication feature that generates consistent vocal profiles from a 30-second audio sample, backed by built-in consent verification protocols.
Security and transparency remain central to the new release, with every generated audio clip featuring SynthID watermarking to ensure that synthetic speech remains detectable and secure against misuse. The platform also incorporates C2PA credentials to protect both developers and professional voice talent from unauthorized replication or identity theft.
Performance and Line-by-Line Direction
Both text-to-speech models offer precise control over performance delivery, allowing users to write custom stage directions or prompt natural script cues for specific emotional tones. The technology supports long-form generation, maintaining consistent voice quality and natural pacing across hours of continuous audio without speaker drift. Additionally, native two-speaker scene staging enables seamless multi-turn conversations from a single script. Microsoft Dismantles EvilTokens Cybercrime Platform That Used AI Chatbots to Breach 12,000 Accounts Globally; Details Here.
The models have achieved top rankings on Hume AI benchmarks, securing the leading position in overall voice design and accent modelling. Developers can access these capabilities immediately through the Gemini API and Google AI Studio, while enterprise integrations roll out across various global partner platforms to accelerate media localization and voice agent deployment.
(The above story first appeared on LatestLY on Sep 23, 2026 11:09 PM IST. For more news and updates on politics, world, sports, entertainment and lifestyle, log on to our website latestly.com).