Google announced Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking on September 15, 2026, a pair of live dialogue models rolling out across the Gemini API, Google AI Studio, Gemini Enterprise, Search Live, Gemini Live, and Google Workspace.
The models were introduced by Tom Ouyang, a principal engineer, and Malini Jaganathan, a member of technical staff, writing on behalf of the Gemini Audio Team. Google describes the pair as its most advanced live dialogue models yet, built for natural conversation, and says gains in intelligence and parallel reasoning make them more intuitive to collaborate with on complex tasks executed by voice. The company positions Gemini 3.8 Live for scale and cost efficiency, combining conversational intelligence with fluid dialogue and visual grounding, while Gemini 3.8 Live Extended Thinking is built for high-complexity tasks that call for increased intelligence and multi-step reasoning. According to Google, the models give developers and enterprises the building blocks for reliable, production-ready voice agents while making conversations with Gemini across the Gemini app, Workspace, and Search more fluid.
Reported Benchmark Results
Google reports that Gemini 3.8 Live Extended Thinking captured the top overall position on Artificial Analysis’ Speech to Speech Quality Index with a score of 82.6, and that it leads in agentic task completion with 68.6% on τ-Voice and 35.1% on Sierra’s τ-Voice-banking benchmark. The company also reports a 97.7% score on Big Bench Audio and says Extended Thinking maintains a highly competitive price point relative to other frontier models. Google says Gemini 3.8 Live placed second in Artificial Analysis’ Speech Agent Arena, citing high preference among users, and describes it as highly cost-effective and built for scale. On ServiceNow’s EVA-Bench, a benchmark for evaluating voice agents, Google says its models push the Pareto Frontier for complex workflows by successfully balancing accuracy with conversational quality; the post notes this evaluation was run on the Live API on Gemini Enterprise Agent Platform.
Real-Time Conversation Capabilities
According to Google, Gemini 3.8 Live processes visual inputs in near real time, adding context the company says produces more helpful responses. The model automatically detects and transitions between 97 supported languages mid-conversation, and it executes tools and API calls in the background while the conversation continues, acknowledging requests and keeping the dialogue going while tasks finish. For tasks that require deeper reasoning, Google says Extended Thinking reasons and speaks simultaneously, using early verbal cues to acknowledge prompts naturally and live progress narration to walk users through multi-step background tasks as they advance, without breaking the conversational flow.
Demonstration videos published with the announcement show Gemini 3.8 Live guiding an employee onboarding session in real time using visual context to answer live questions, playing chess in near real time, building complete business plans and custom marketing toolkits through natural speech, and powering step-by-step troubleshooting help inside Search Live. Extended Thinking demonstrations show the model turning raw sketches and near-real-time voice feedback into functional React components, coordinating multi-step restaurant reservations through asynchronous function calls without interrupting the conversation, and working across Docs Live, Gmail Live, and Keep Live in Google Workspace.
Live API and Developer Ecosystem
Google names Agora, Fishjam, LiveKit, Pipecat, Vercel, and Vision Agents as developer platforms that use the Gemini Live API to let developers build and deploy voice-driven interfaces while the platforms manage the underlying real-time media streaming infrastructure. The company says it is also partnering with Salesforce, Genspark, and Lumeris on the new models, pointing to their interest in the models’ latency, fluidity, and tool-calling capabilities.
The official Live API documentation describes the interface as enabling low-latency, real-time voice and vision interactions with Gemini, processing continuous streams of audio, images, and text to deliver immediate spoken responses over a stateful WebSocket connection. Documented inputs are raw 16-bit PCM audio at 16kHz, JPEG images at up to one frame per second, and text, with audio output as raw 16-bit PCM at 24kHz. Developers can choose a server-to-server implementation, in which a backend forwards client stream data to the Live API, or a client-to-server implementation, in which frontend code connects directly over WebSockets. The documentation lists use cases spanning retail shopping assistants and support agents, interactive game characters, voice- and video-enabled experiences in robotics, smart glasses, and vehicles, health companions, financial advisory tools, education mentors, real-time translation, and live transcription and captioning.
Google states that all audio generated by its AI products is watermarked with SynthID, an imperceptible watermark woven directly into the audio output so that AI-generated content remains detectable to help prevent misinformation. The post points readers to the model card for details on the company’s safety and responsibility approach.
Availability and Rollout
Both models began rolling out on September 15. For developers, they are available in the Gemini API and Google AI Studio, and for enterprises they are in private preview in Gemini Enterprise. Gemini 3.8 Live is available to everyone in Search Live, while Extended Thinking is available in Gemini Live, in Docs for Google AI Pro and Ultra subscribers through Workspace, and in Gmail and Keep for all Google AI subscribers. Google says Gemini Enterprise for Customer Experience support for both models is coming soon, and that Extended Thinking is coming soon to Google Workspace business customers.
Credit: Source link


























