
Google DeepMind launched Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking on September 15, 2026 — two voice-first models that reason and speak in parallel, and immediately claimed the top spot on Artificial Analysis's Speech-to-Speech Quality Index with a score of 82.6.
The release marks Google's most serious push yet into real-time voice AI, a category that has been dominated by OpenAI's GPT-Live models and xAI's Grok Voice. Both new models are available now through the Gemini API and Google AI Studio.
What's new
Gemini 3.8 Live is built for low-latency, natural conversation. It can automatically switch between 97 languages mid-conversation, process live visual input from a camera feed, and execute tools and API calls in the background while the dialogue keeps flowing — no awkward pauses while the model thinks.
The Extended Thinking variant adds deeper reasoning for complex, multi-step tasks. It can narrate its progress verbally while working through problems in the background, using verbal acknowledgment cues to let the user know it is still engaged.
Benchmark performance
| Benchmark | Gemini 3.8 Live ET | GPT-Live-1-Astra | Grok Voice Think Fast 2.0 |
|---|---|---|---|
| S2S Quality Index (Artificial Analysis) | 82.6 (#1) | Below 82.6 | Below 82.6 |
| τ-Voice agentic task completion | 68.6% | Not reported | Not reported |
| Sierra τ-Voice-banking | 35.1% | Not reported | Not reported |
| Big Bench Audio reasoning | 97.7% | Not reported | Not reported |
Source: Google DeepMind announcement, September 15, 2026. Competitor scores on τ-Voice and Big Bench Audio were not published alongside Google's results.
The 82.6 score on Artificial Analysis's S2S Quality Index puts Gemini 3.8 Live Extended Thinking ahead of both GPT-Live-1-Astra and Grok Voice Think Fast 2.0, according to Google. The model also scored 97.7% on Big Bench Audio, a benchmark that tests audio reasoning rather than simple transcription accuracy.
On agentic voice tasks, the Extended Thinking model completed 68.6% of tasks on the τ-Voice benchmark and 35.1% on Sierra's more rigorous τ-Voice-banking test, which simulates customer service interactions.
Pricing
| Usage | Price |
|---|---|
| Audio input | $0.005 per minute |
| Audio output | $0.018 per minute |
Source: Google AI Studio pricing, September 2026
Why it matters
Voice AI has been the missing piece in the assistant race. Text models have become extraordinarily capable, but real-time spoken interaction remains hard — latency, naturalness, and the ability to handle interruptions all degrade quickly when a model has to think out loud. Google's claim of #1 on the S2S Quality Index matters because that benchmark is designed to measure exactly these conversational qualities, not just raw intelligence.
The parallel reasoning architecture is the technically interesting part. Rather than the traditional "think, then speak" pipeline, these models interleave reasoning and speech — they can start talking while still working through a problem, and they can execute API calls in the background without stopping the conversation. This is closer to how humans actually converse, and it could make voice agents feel far less robotic in production settings like customer support.
That said, the benchmark picture is incomplete. Google published its own scores on τ-Voice and Big Bench Audio but did not include comparable numbers for GPT-Live-1-Astra or Grok Voice on those same tests. The S2S Quality Index ranking is third-party, which lends credibility, but the agentic task scores are self-reported. Independent verification will be needed before the #1 claim holds up across the board.
The pricing is aggressive. At $0.005 per minute for input and $0.018 for output, a typical 5-minute voice conversation costs roughly $0.06 — well below what many enterprise voice platforms charge. If the quality holds up in real-world deployments, this could pressure OpenAI and xAI to cut voice API prices.
Google has historically struggled to monetize its AI research — DeepMind produces world-class models, but product adoption has lagged behind OpenAI and Anthropic. Gemini 3.8 Live could change that if it becomes the default voice layer for Google Cloud's contact center and enterprise assistant offerings. The 97-language support is a particular advantage for global enterprises that OpenAI and xAI cannot easily match.
What to watch
- Independent third-party benchmarks on τ-Voice and Big Bench Audio to verify Google's self-reported scores
- Enterprise adoption announcements — Google Cloud contact center integrations would be the strongest signal
- Whether OpenAI responds with a GPT-Live update or price cut before the end of September
- Real-world latency measurements from developers using the API in production
No comments yet