Best Voice AI Models 2026: GPT-Live-1 vs Gemini 3.8 Live vs Grok Voice
Quick verdict
Google's Gemini 3.8 Live Extended Thinking took the top of Artificial Analysis's Speech-to-Speech Index at 82.6 on September 15, 2026, hours after OpenAI's GPT-Live-1 had claimed first place at 81.5 — with xAI's Grok Voice Think Fast 2.0 close behind at 81.3. The 0.3-point spread is noise; the real differences are architecture (full-duplex vs turn-based), component scores on agentic tasks, and price per hour of audio, where Gemini 3.8 Live costs $0.84 versus $5.83 for GPT-Live-1's Astra configuration.
Two Days, Two Number Ones
The voice model leaderboard changed hands twice inside a single day, and the sequence is more instructive than either announcement.
On the morning of September 15, 2026, Artificial Analysis published its Speech-to-Speech Index with OpenAI's GPT-Live-1 on top at 81.5. Hours later, Google DeepMind released Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking — and the Extended Thinking variant debuted in first place at 82.6, pushing GPT-Live-1 to second and xAI's Grok Voice Think Fast 2.0 High to third at 81.3.
The board as captured on September 16:
| Rank | Model | Index score |
|---|---|---|
| 1 | Gemini 3.8 Live Extended Thinking | 82.6 |
| 2 | GPT-Live-1 (Astra backend, medium effort) | 81.5 |
| 3 | Grok Voice Think Fast 2.0 High | 81.3 |
| 4 | GPT-Live-1 (Sol backend, low effort) | 80.1 |
| 5 | Gemini 3.8 Live (standard) | 76.0 |
| 6 | GPT-Realtime-2.1 High | 73.9 |
| 7 | Gemini 3.1 Flash Live High | 71.5 |
Three things fall out of that list, and none of them is "Google won."
First, the headline hides the structure. The model in the matchup's title — standard Gemini 3.8 Live — sits fifth at 76.0. Google's first place belongs to the Extended Thinking variant, which is a different product at a different price with a different rollout.
Second, GPT-Live-1 appears twice, because Artificial Analysis scores the whole system rather than the voice front end. The 1.4-point gap between its Astra and Sol configurations is nearly five times larger than the 0.2 points separating first from second place.
Third, the 0.3 points between first and third are noise. A composite index is a weighted average; a gap that small tells you the three models are close, not which one is better.
Full-Duplex vs Turn-Based — and Why It Stopped Deciding Things
For about two months, the architecture question settled the comparison. GPT-Live-1 is genuinely full-duplex: it processes input and output audio continuously inside a single model rather than chaining speech-to-text into a language model into text-to-speech. That eliminates information loss between stages and produces the behavior people actually notice — you can cut in mid-answer and it adapts; it does not mistake a pause or background noise for the end of your turn; it emits backchannels like "mhmm" and stays quiet until called on.
Gemini 3.8 Live is turn-based. Under the old framing, that ended the argument. It no longer does, because Google shipped the thing full-duplex was compensating for: a model that holds a conversation while a multi-step task runs in the background, and a variant that reasons out loud while working — acknowledging a request, then narrating progress as it executes tools and API calls.
Both models refuse to hang up on you while they think. GPT-Live-1 does it by never stopping the audio stream. Google does it by talking over the work. The architectural difference is real; it is just no longer the axis that decides the purchase.
Read the Components, Not the Composite
Composites hide arguments. The Speech-to-Speech Index averages speech reasoning, agentic performance, arena preference, and task success rate — and these models are not close on all four.
| Benchmark | Gemini 3.8 Live ET | GPT-Live-1 | Grok Voice Think Fast 2.0 |
|---|---|---|---|
| τ-Voice (agentic) | 68.6% | 67.9% | 56.5% |
| τ³-Banking (Sierra) | 35.1% | 32.0% | 16.5% |
| Big Bench Audio (reasoning) | 97.7% | — | 97.2% |
| Time to first audio | 1.18s / 1.35s (ET) | — | 0.70s (fastest tested) |
Two readings matter here. Google has taken the agentic category xAI was leading — on τ-Voice, Grok's 56.5% trails by roughly 12 points. And Grok remains the fastest to first audio, which is the difference a caller actually feels. Speech reasoning, meanwhile, is effectively level: 97.7% versus 97.2%, with Qwen Audio 3.0 Realtime Plus ahead of both at 99.2%.
Gemini 3.8 Live also carries some product details worth knowing: mid-conversation switching across 97 languages, background tool and API execution that continues while the conversation goes on, near-real-time visual input, and SynthID watermarks on generated audio. The API model IDs are gemini-3.8-live and gemini-3.8-live-extended-thinking; the standard model is available in the Gemini API and AI Studio, and Extended Thinking reaches Gemini Live, Docs, Gmail, and Keep for AI Pro and Ultra subscribers.
The Real Change: Pricing Moved to Hours
The most consequential number in this release is not a score. It is the unit.
Artificial Analysis prices these models by cost per hour of input audio — roughly 25 tokens per second of audio. That reframing matters because it matches how voice agents are actually bought: a support seat talks for hours.
| Model | Cost per hour of input audio |
|---|---|
| Gemini 3.8 Live (standard) | $0.84 |
| Gemini 3.8 Live Extended Thinking | $3.50 |
| Grok Voice Think Fast 2.0 | $4.80 |
| GPT-Live-1 (Astra configuration) | $5.83 |
| GPT-Realtime-2.1 High | $10.75 |
Run the seat math. A support agent handling eight hours of conversation per day pays under $7 daily on Gemini 3.8 Live, about $47 on GPT-Live-1's Astra configuration, and roughly $86 on GPT-Realtime-2.1 High. Multiply by seats and by days.
Google's standard model is the cheapest on the index and roughly six times cheaper than GPT-Live-1's Astra configuration — about half the price of its own previous generation.
One caveat that changes the arithmetic. OpenAI's launch pricing for GPT-Live-1 was $0.05 per minute for the voice front end — about $3 per hour — which is less than the $5.83 Artificial Analysis lists for the same model. The gap is not a contradiction: GPT-Live-1 delegates hard queries to a frontier text model (GPT-5.5 and GPT-6 Astra are both documented backends, with reasoning-effort selectors exposed), and those tokens bill separately. The cheaper headline price covers the conversation layer, not the thinking.
Any serious cost model has to include both layers. Google's integrated design makes that simpler to estimate; OpenAI's separated design makes it more tunable — you can route easy turns cheaply and escalate only the hard ones.
How to Choose
| Your situation | Recommendation |
|---|---|
| High-volume inbound calls, mostly simple | Gemini 3.8 Live — $0.84/hour is hard to argue with when agentic depth is not the bottleneck |
| Agentic voice workflows (booking, transactions, multi-step tool use) | Gemini 3.8 Live Extended Thinking or GPT-Live-1 — the τ-Voice and τ³-Banking gaps are where they earn the premium |
| Conversational naturalness is the product | GPT-Live-1 — full-duplex turn-taking remains the benchmark for interruption handling |
| Latency-critical real-time interaction | Test Grok — 0.70s time to first audio leads the field |
| Mixed workload | Route simple turns to a cheap model and escalate task-heavy turns, as long as the handoff preserves conversation state |
| Voice agent on your own infrastructure | None of these are self-hostable at this tier; plan for API dependency |
One methodological warning applies throughout. Vendor figures in this category are not durable: when xAI announced Grok Voice Think Fast 2.0 in late July 2026, the reported index figure was 82.9, ahead of GPT-Realtime-2.1 at 79.1. The same model now reads 81.3 on the public board. That is not an accusation of dishonesty — composites get revised and a July number measured against a July field is simply old. But it means the only figure worth planning against is one you read yourself, with a date attached.
Summary
The voice model race is now close enough that composite index positions are marketing, not purchasing guidance. What actually separates these three is narrower and more useful: GPT-Live-1 for conversational naturalness and delegation depth, Gemini 3.8 Live for agentic tasks and price, Grok Voice for raw latency — with Google's own standard and Extended Thinking tiers sitting on opposite ends of the value curve.
The pricing shift is the bigger story. When the unit moves from tokens to hours of audio, voice agents become comparable to human line items — a few dollars per seat-day against the cost of the seat itself. That is the comparison enterprise buyers will make, and it is why $0.84 per hour matters more than 82.6 versus 81.5.
For the API layer OpenAI is building underneath, see our Agents API guide; for text-side comparisons, GPT-5.6 Sol and Gemini.
Related Articles
Keep reading