Gemini 3.8 Live:交錯推理(interleaved reasoning)與非同步工具呼叫如何改變 Google 的語音代理

编辑者: Svitlana Velhush

2026 年九月 15 日,Google 推出了兩款新模型——Gemini 3.8 Live 和 Gemini 3.8 Live Extended Thinking。它們被定位為最先進的即時對話模型,針對低延遲和即時性進行了優化。

與 Gemini 3.1 Flash Live 等前代產品的主要區別在於對交錯推理(interleaved reasoning)和非同步函數呼叫的支援。該模型能夠在推理的同時生成語音,而工具則在後台執行,不會中斷對話。

這是透過原生音訊架構實現的,輸入和輸出均為即時音訊與文字。支援近乎即時(near real-time)的視覺上下文、在 97 種語言之間自動切換以及會話內容更新。價格為每分鐘音訊輸入 $0.005,每分鐘輸出 $0.018。

Extended Thinking 在基準測試中處於領先地位:在 Artificial Analysis Speech to Speech Quality Index 獲得 82.6 分,在 τ-Voice 獲得 68.6% 分,在 Sierra’s τ-Voice-banking 獲得 35.1% 分,以及在 Big Bench Audio 獲得 97.7% 分。基礎版本在 Speech Agent Arena 中排名第二,並在 EVA-Bench 上表現強勁。

基準測試方法側重於代理任務完成度(agentic task completion)和語音品質,但並不總是揭露留出(held-out)測試或與零樣本(zero-shot)場景比較的細節。這讓人對其在特定企業任務之外的實際泛化能力存疑。

在產業格局中,這與 OpenAI 的 GPT-Live 和 xAI 的 Grok Voice Think Fast 的方法有所不同。Google 強調任務的並行執行和視覺基礎(visual grounding),而不僅僅是回應速度。先前的模型通常需要暫停進行推理,這破壞了對話的自然流暢度。

新功能為生產級語音代理鋪平了道路,模型可以在保持對話流暢的同時執行複雜的多步驟操作。這在與 Search Live、Gemini Live 和 Workspace 的整合中尤為明顯。

目前尚不清楚在獨立驗證以及真實環境噪音或口音條件下,這些結果的穩定性如何。社群可能會測試邊緣案例(edge-cases)中的延遲,並將其與開源替代方案進行比較。

關鍵的轉變是從反應式語音介面轉向真正的代理式(truly agentic)介面,其中推理過程不會干擾自然的交流。

3 浏览量

來源

  • Google launches Gemini 3.8 Live models

留言

阅读更多关于此主题的文章:

Replying to @agentcommunity_

The Navier-Stokes dispute raises unsettled questions on AI research credit. Mathematicians Buckmaster and Alpöge reportedly made progress with AI assistance; reports claim OpenAI then prompted models on the same direction and possibly dropped Alpöge from authorship. OpenAI denies

Emad
Emad
@EMostaque

Wow, Navier-Stokes drama This statement is worth reading in full from Tristan Buckmaster discussing his work with @__alpoge__ and OpenAI’s upcoming Condition C/D result (!) Crazy cims.nyu.edu/~tristanb/stat… mastodon.social/@tristanbuckma…

Image
Reply
发现错误或不准确的地方吗?我们会尽快处理您的评论。