Artificial intelligence
Google Won the Voice Benchmark by 1.1 Points. It Undercut the Price by Seven Times
Gemini 3.8 Live Extended Thinking takes the top spot on Artificial Analysis' speech quality index by 1.1 points over OpenAI's GPT-Live-1 Astra. The interesting number is the price: roughly a seventh of what the model it beat costs per hour.
MAI
Google released two conversational models on 15 September: Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking. The claim it led with is the top position on Artificial Analysis' Speech to Speech Quality Index, where Extended Thinking scores 82.6.
The number sitting underneath that one is more interesting. OpenAI's GPT-Live-1 Astra, the model it displaced, scores 81.5. Grok Voice Think Fast 2.0 scores 81.3. Three labs are now within 1.3 points of each other on the metric that is supposed to separate them. On a benchmark with that spread, first place is a press release, not a moat.
Where the actual gap is
The separation is in price, and it is not marginal.
| Model | Speech to Speech Quality Index | τ-Voice | τ-Voice-banking | Reported cost per hour |
|---|---|---|---|---|
| Gemini 3.8 Live Extended Thinking | 82.6 | 68.6% | 35.1% | $3.50 |
| Gemini 3.8 Live | — (2nd, Speech Agent Arena) | — | — | $0.84 |
| GPT-Live-1 Astra (Medium) | 81.5 | 67.9% | 32.0% | $5.83 |
| Grok Voice Think Fast 2.0 (High) | 81.3 | 56.5% | 16.5%* | $4.80 |
\xAI-Realtime on the banking benchmark. Hourly figures are third-party calculations from published rates, not Google's own comparison.*
Google's developer post prices the Live API at $3 per million input tokens and $12 per million output tokens, or $0.005 per minute of audio in and $0.018 per minute of audio out. Worked through to a billable hour of conversation, that puts the cheaper model at roughly a seventh of what OpenAI charges for a model it trails by 1.1 points.
That is the argument Google is actually making, and it is aimed squarely at the people building call centres, drive-throughs and support lines rather than at the leaderboard. A contact centre running ten thousand hours a month does not care about the first decimal place of a quality index. It cares that the same workload costs $8,400 instead of $58,300.
The benchmarks themselves are third-party — Artificial Analysis for the quality index, Sierra's τ³-Banking leaderboard for the agentic scores — which is worth more than a vendor's internal eval. Google still chose which of them to put on the slide.
Reasoning out loud
The technical claim behind Extended Thinking is that it reasons and speaks at the same time. Google describes the model as using "early verbal cues to acknowledge prompts naturally" and giving "live progress narration to walk users through multi-step background tasks."
That is a direct answer to the tradeoff that has defined voice agents since they became usable. A model that thinks before it talks is accurate and sounds broken; a model that talks immediately is fluent and shallow. Every deployment so far has papered over the gap with filler audio — the synthetic "let me check that for you" while a slower model works in the background. Google's pitch is that the narration is no longer filler, because the reasoning is genuinely running while the speech is being produced.
This is the same instinct visible in OpenAI's Astra, which we reviewed yesterday: it gets to the top of the intelligence rankings on roughly a third of the output tokens of its nearest rival. Both labs have concluded that the frontier in conversation is no longer raw capability but how much work you can avoid doing before the user hears something.
The number Google did not lead with
Extended Thinking completes 35.1 per cent of tasks on Sierra's τ-Voice-banking benchmark. That is the best score anyone has posted. It is also a model failing roughly two out of three realistic banking requests.
τ-Voice-banking is deliberately hard — multi-turn customer service with tool calls, policy constraints and a caller who does not read from a script. The general τ-Voice number is nearly twice as high, at 68.6 per cent, which tells you how much of the difficulty comes from the regulated, tool-dependent part rather than from the talking.
The honest reading of this release is that conversational quality is close to solved and agentic voice is not. You can now buy a model that sounds human for under a dollar an hour. You cannot yet buy one that can be left alone with a customer's account.
Where it ships
Both models are available today through the Gemini API and Google AI Studio, with integrations for Pipecat, LiveKit, LangChain, Agora, Fishjam, Vercel and Vision Agents. Gemini 3.8 Live is in Search Live for everyone; Extended Thinking reaches Gemini Live for Pro and Ultra subscribers and appears in Gmail and Keep. Enterprise access is a private preview. The Live models cover 97 or more languages.
Google also shipped Gemini 3.5 Transcribe alongside them: 85 or more languages, a 4.0 per cent word error rate streaming and 2.6 per cent non-streaming. It got no headline, which is roughly correct, and it is the piece most existing production systems will swap in first.
Sources: Google: Introducing Gemini 3.8 Live and 3.8 Live Extended Thinking · Google: Build real-time voice applications with Gemini 3.8 Live and 3.5 Transcribe · Unite.ai: Google Launches Gemini 3.8 Live and Extended Thinking Voice Models · OfficeChai: Google Releases Gemini 3.8 Live-Extended Conversational Model · Thurrott: Google Announces Gemini 3.8 Live and 3.8 Live Extended Thinking