On September 15, 2026, Google introduced Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, two live dialogue models designed for real-time voice agents. Extended Thinking tops Artificial Analysisβ Speech-to-Speech Quality Index at 82.6 and supports multi-step reasoning while continuing to speak. Both models handle visual inputs in near real time, switch among 97 languages mid-conversation, and execute tool calls asynchronously. Developers can access them via the Gemini API and Google AI Studio; enterprise private preview is available in Gemini Enterprise.
Voice agents moved from demo to production infrastructure in a single model drop. Gemini 3.8 Live and its Extended Thinking variant give technology teams lower-latency, higher-reasoning building blocks for customer support, internal copilots, and multimodal workflows that previously required stitching multiple systems together.
- What changed? Two new live audio models with parallel reasoning, visual grounding, 97-language support, and asynchronous tool calling.
- When? Announced and rolling out September 15, 2026 for developers; enterprise private preview concurrent.
- Who is affected? Product, engineering, and AI platform teams building voice or multimodal agents.
- What to do this week? Benchmark latency and task-completion rates against current voice stacks, test tool-calling patterns in the Live API, and map Workspace Live features for internal users.
What did Google release?
Gemini 3.8 Live is positioned as the scalable, cost-efficient model for fluid dialogue and visual grounding. Gemini 3.8 Live Extended Thinking adds higher intelligence and multi-step reasoning while maintaining conversational flowβusing early verbal acknowledgments and live progress narration so users are not left waiting in silence. Both accept text, images, audio, and video; both output text and audio; both support function calling (asynchronous on Extended Thinking). Context windows are listed at 131,072 input tokens and 65,536 output tokens. All generated audio carries a SynthID watermark.
Google reported that Extended Thinking scored 82.6 on Artificial Analysisβ Speech-to-Speech Quality Index (first place in the cited ranking), 68.6% on Ο-Voice agentic task completion, and 97.7% on Big Bench Audio. Gemini 3.8 Live placed second in the Speech Agent Arena user preference ranking.
How do the models change agent architecture?
Previous generations often forced a trade-off between low latency and deep reasoning. Extended Thinking is designed to reason in the background while continuing to stream audio, reducing the awkward pauses that break user trust. Asynchronous tool calls allow the model to keep talking while APIs complete. Near-real-time visual grounding lets an agent reference what a user is looking at or showing on camera. For technology teams, this means fewer custom orchestration layers and a clearer path to production-grade voice agents inside customer-experience and internal-tool workflows.
Where can teams start using them?
Developers have access through the Gemini API and Google AI Studio. Enterprises can join private preview in Gemini Enterprise, with broader availability planned for Gemini Enterprise for Customer Experience. End-user surfaces include Search Live, Gemini Live, and Workspace features (Docs Live for Pro/Ultra subscribers; Gmail and Keep Live for Google AI subscribers). Google also listed partner platforms already integrating the Live API, including LiveKit, LangChain, Pipecat, and others.
What should technology teams do this week?
Run side-by-side latency and task-completion tests against your current voice or multimodal stack. Prototype one high-value internal workflow (onboarding, support triage, or document assistance) using the Live APIβs asynchronous tool pattern. Confirm data-handling and watermarking requirements with legal and security. Brief product owners on the Workspace Live rollout so internal early adopters can be identified without creating shadow-IT risk.
What to watch next
Watch general-availability timing for Gemini Enterprise for Customer Experience, pricing clarity for high-volume audio sessions, and any additional language or modality expansions. Also track how competitors respond with their own full-duplex reasoning models; the quality-and-cost frontier is moving quickly.
Frequently Asked Questions
Are the models available today?
Yes for developers via the Gemini API and Google AI Studio. Enterprise private preview is open in Gemini Enterprise.
What is the difference between Live and Live Extended Thinking?
Live prioritizes scale and cost efficiency with strong conversational quality. Extended Thinking adds deeper multi-step reasoning while still speaking in real time.
Do the models support tool calling?
Yes. Both support function calling; Extended Thinking emphasizes asynchronous execution so conversation continues while tools run.
Is audio watermarked?
Yes. Google states that all audio generated by its AI products carries a SynthID watermark.
Which languages are supported?
Google lists automatic detection and mid-conversation switching among 97 languages.
Son GΓΌncelleme / Last Updated: September 17, 2026
Related: Agentic AI 2026 Β· Technology News Β· Technology Hub
Discover more from Kurums | Business Intelligence
Subscribe to get the latest posts sent to your email.