Gemini 3.8 Live and Extended Thinking models launch with enhanced dialogue capabilities

The Gemini Audio Team has unveiled its latest advancements in AI dialogue technology with the launch of Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking. These models are designed to offer more intuitive interactions and support complex tasks through enhanced intelligence and parallel reasoning capabilities.

These models aim to make voice interactions seamless, allowing users to engage in conversations without interruptions while the system handles complex reasoning, real-time visual context integration, and background task execution. Available today through the Gemini API, Google Workspace, and the Gemini app, these features are set to revolutionize voice-based interactions.

Google’s new AI models enhance the natural feel of device communications by managing interruptions, language transitions, and offering explanations of their processes. Whether tackling intricate problems or casual discussions, users will experience AI that listens and responds thoughtfully, simulating a genuine conversational partner.

For developers and enterprises, Gemini’s models provide a foundation for building dependable voice agents. The models integrate fluently with tools like the Gemini app, Google Workspace, and Search, facilitating complex tasks through voice commands.

Gemini 3.8 Live Extended Thinking excels in enterprise-grade task completion, ranking first on Artificial Analysis’ Speech to Speech Quality Index with a score of 82.6. It leads in agentic task completion and demonstrates strong reasoning capabilities on Big Bench Audio. Meanwhile, Gemini 3.8 Live is preferred for its efficiency and affordability, scoring second in the Speech Agent Arena.

Performance assessments on ServiceNow’s EVA-Bench indicate that these models successfully balance accuracy and conversational quality, advancing the capabilities of voice agents for complex workflows.

Gemini 3.8 Live processes visual inputs in near real-time, enhancing responses through contextual understanding. It supports 97 languages, switching seamlessly during interactions, and executes background tasks while maintaining active conversations.

The 3.8 Live Extended Thinking variant addresses tasks requiring deep reasoning by providing simultaneous reasoning and speech. It enhances complex workflows without disrupting the conversational flow, using natural verbal cues and progress narration.

These innovations extend across Google Workspace and Search, offering more intuitive and collaborative experiences, especially for complex tasks. Developers can leverage the Gemini Live API on platforms like Agora and Vercel, focusing on user experience while the system manages real-time media streaming infrastructure.

Collaborations with companies such as Salesforce and Lumeris highlight the models’ capabilities in latency, fluidity, and tool-calling. All audio outputs are safeguarded with SynthID, an imperceptible watermark to ensure AI-generated content remains identifiable, supporting the prevention of misinformation.

Gemini 3.8 Live and 3.8 Live Extended Thinking are available starting today. Interested users can sign up for updates and offers through newsletters, with information managed under Google’s privacy policy.