Gemini 3.8 Live: How interleaved reasoning and asynchronous tool calls are changing Google's voice agents

Edited by: Svitlana Velhush

On September 15, 2026, Google introduced two new models—Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking. They are positioned as the most advanced models for live dialogue, optimized for low latency and real-time interaction.

The main difference from predecessors, such as Gemini 3.1 Flash Live, lies in the support for interleaved reasoning and asynchronous function calls. The model can reason and generate speech simultaneously, while tools execute in the background without interrupting the conversation.

This is achieved through a native audio architecture where input and output are audio and text in real time. It supports visual context in near real-time, automatic switching between 97 languages, and session content updates. The price is $0.005 per minute of audio input and $0.018 per minute of output.

Extended Thinking leads in benchmarks: 82.6 points on the Artificial Analysis Speech to Speech Quality Index, 68.6% on τ-Voice, 35.1% on Sierra’s τ-Voice-banking, and 97.7% on Big Bench Audio. The base version takes second place in the Speech Agent Arena and shows strong results on EVA-Bench.

The benchmark methodology focuses on agentic task completion and speech quality, but it does not always disclose details of held-out tests or comparisons with zero-shot scenarios. This leaves questions about real generalization beyond specific enterprise tasks.

In the landscape, this differs from the approaches of OpenAI with GPT-Live and xAI with Grok Voice Think Fast. Google emphasizes parallel task execution and visual grounding, rather than just response speed. Previous models often required pauses for reasoning, which disrupted the natural flow of dialogue.

New capabilities pave the way for production-grade voice agents, where the model can perform complex multi-step operations while maintaining the conversational flow. This is especially noticeable in integrations with Search Live, Gemini Live, and Workspace.

It remains unclear how robust the results are under independent verification and in conditions of real-world noise or accents. The community will likely test latency in edge cases and compare it with open-source alternatives.

The key shift is the transition from reactive voice interfaces to truly agentic ones, where reasoning does not interfere with natural communication.

3 Views

Sources

  • Google launches Gemini 3.8 Live models

Comments

Read more articles on this topic:

Replying to @agentcommunity_

The Navier-Stokes dispute raises unsettled questions on AI research credit. Mathematicians Buckmaster and Alpöge reportedly made progress with AI assistance; reports claim OpenAI then prompted models on the same direction and possibly dropped Alpöge from authorship. OpenAI denies

Emad
Emad
@EMostaque

Wow, Navier-Stokes drama This statement is worth reading in full from Tristan Buckmaster discussing his work with @__alpoge__ and OpenAI’s upcoming Condition C/D result (!) Crazy cims.nyu.edu/~tristanb/stat… mastodon.social/@tristanbuckma…

Image
Reply
Did you find an error or inaccuracy?We will consider your comments as soon as possible.