GPT-6 Sol and GPT-6 Luna Review: Half the Price, and the Tier Below (2026)
Quick verdict
OpenAI released GPT-6 Sol and GPT-6 Luna on September 22, 2026, priced 50% below GPT-5.6 promo rates: Sol at $2/$10 per million tokens and Luna at $0.10/$0.50, against Astra's unchanged $10/$50. All three share a 1.05M-token context window; they differ in capability and cost. Sol is the general-purpose agentic tier at $0.27 per AutomationBench task; Luna is a high-volume worker that drops to 20.7% on AutomationBench and carries a 7.6% factual error rate.
What Shipped
OpenAI released GPT-6 Sol and GPT-6 Luna on September 22, 2026 — hours after Anthropic released Claude Opus 5.5. Both are built on the advances behind GPT-6 Astra and priced 50% below GPT-5.6 promotional rates, with the cut attributed to improvements in caching and inference that OpenAI says it is passing through.
| Model | Input / 1M | Output / 1M | Cached input / 1M |
|---|---|---|---|
| GPT-6 Astra | $10.00 | $50.00 | $1.00 |
| GPT-6 Sol | $2.00 | $10.00 | $0.20 |
| GPT-6 Luna | $0.10 | $0.50 | $0.01 |
Context is identical across all three: 1,050,000 tokens with 128,000 max output, text and image input, text output, reasoning effort from none to max with medium as the default. Knowledge cutoffs differ — Sol at April 20, 2026, Luna at May 18, 2026.
Availability: the API (gpt-6-sol, gpt-6-luna), plus ChatGPT Work and Codex for Plus, Pro, Business, Enterprise, and Edu. Free and Go users get Luna in the desktop app, which puts the cheapest tier in front of everyone rather than behind a subscription.
The Number OpenAI Did Not Lead With
The launch framing was price. The more consequential figure is this: Luna scores 66.6% on DeepSWE v1.1 against Sol's 68.8% — a 2.2-point gap, while Luna costs one twentieth of Sol and one hundredth of Astra on input.
That is an extraordinary value line. But the rest of the benchmark table explains why it comes with a catch:
| Benchmark | Luna | Sol | Astra |
|---|---|---|---|
| DeepSWE v1.1 | 66.6% | 68.8% | 74.1% |
| FrontierCode 1.1 | 42.4% | 49.3% | 53.3% |
| AutomationBench | 20.7% | 33.2% | 41.4% |
| Agents' Last Exam | 50.9% | 56.4% | 59.3% |
| OSWorld 2.0 | 52.7% | 64.4% | 73.5% |
| Factual error rate | 7.6% | 4.6% | 3.9% |
Read those rows together and the tiers describe something specific. Luna writes competent code when the task arrives in one piece, and loses its way when it has to decide what to do next. AutomationBench drops 12.5 points from Sol; OSWorld 2.0 — computer use — drops 11.7. Its factual error rate is nearly double Astra's.
Luna is a worker, not a planner. That is a defensible product decision, and saying it plainly is more useful than framing the whole lineup as one capability curve with different prices.
Sol Is the Interesting Tier
Sol is where the launch actually competes. OpenAI published a cost-per-task figure — the right way to report agentic performance, because per-million-token pricing tells you little about what an agent run costs:
- AutomationBench: 33.2% at $0.27 per task, which OpenAI says beats Claude Opus 5 at max effort at 9% of its cost per task, and exceeds low-effort GPT-6 Astra (30.3%) and Claude Fable 5.1 with Opus 5 fallback (31.4%)
- Agents' Last Exam: 56.4% at max effort, above Claude Opus 5's best score at 60% lower cost per task
- DeepSWE v1.1: 68.8% at max effort, within 1.1 points of Claude Fable 5's best (69.9% at xhigh) at roughly 80% lower cost per task
- FrontierCode: substantial improvement over GPT-5.6 Sol, matching Claude Fable 5.1 at xhigh effort for much less
One disclosure worth keeping: OpenAI notes the Fable 5.1 comparison understates that model's real cost, because it omits fallback spending that occurred on roughly 40% of tasks. That is the kind of caveat that makes the rest of the numbers easier to trust.
On internal factuality evaluation, OpenAI reports Sol makes about half as many mistakes as its predecessor, and Luna at higher effort matches GPT-5.6 Sol at roughly a hundredth of the cost.
What Independent Testing Says
Artificial Analysis measured both models as cost halved with the Intelligence Index on par with GPT-5.6 — not a capability jump, a price move. It also reported sharp reductions in hallucination rate: GPT-6 Sol cutting its rate from 92% to 60% and Luna from 93% to 77% on its measurement. Treat the absolute levels of that metric with caution; the direction is the signal.
For market context, the analytics firm Ramp reported that Astra held the largest share of any single AI model by spend the previous week — nearly 19%, ahead of Claude Opus 5 at 17%.
Caching: Read the Fine Print
The 90% cached-input discount is the mechanism behind the price cut, and it has two conditions:
- Reads are cheap, writes are not. Cached input runs at 10% of the uncached rate ($0.20/M on Sol, $0.01 on Luna), but cache writes bill at 1.25 times the uncached input rate.
- Reuse must be prompt. The improved caching system grants the discount for eligible shared prefixes reused within a 30-minute window.
For an agent with a stable system prompt and tool definitions, that is close to free. For workloads that rebuild context every call, expect the effective price to sit well above the headline.
Who Should Use Which Tier
| Workload | Tier | Why |
|---|---|---|
| Classification, extraction, summarization, bulk transforms | Luna | Single-pass tasks where the cost advantage is real |
| First-pass code generation with a reviewer downstream | Luna | 66.6% on DeepSWE is strong for $0.10 input |
| Agentic workflows with tool calls | Sol | The AutomationBench and OSWorld gaps make Luna a false economy here |
| Computer use, browser automation, long-horizon tasks | Astra | 73.5% on OSWorld 2.0 versus Sol's 64.4% is the widest gap in the table |
| Anything customer-facing without human review | Astra | A 7.6% factual error rate on Luna is not shippable |
| Cost-sensitive bulk work needing the best per-task economics | Sol | $0.27 per AutomationBench task is the number to model against |
Summary
GPT-6 Sol and Luna are a price move with a genuine capability step in the middle tier. Sol is the model to evaluate — near-frontier coding results at roughly a fifth of Astra's price, with a published cost-per-task figure that lets you model a real budget. Luna is a volume workhorse, excellent value on single-pass tasks and unsuitable for anything that requires deciding what to do next.
The strategic read is that OpenAI now sells a three-tier ladder sharing one context window, which forces the comparison onto capability per dollar rather than capability alone. Released hours after Claude Opus 5.5, with both cutting prices, the competitive axis this quarter is clearly cost — and the burden of proof sits with the vendors that still charge $50 per million output tokens.
For the flagship behind these models, see our GPT-6 Astra review; for the previous generation, GPT-5.6 Sol, Terra and Luna; for the same-day competitor, Claude Opus 5.5.
Related Articles
Frequently Asked Questions
How much do GPT-6 Sol and Luna cost?+
What is the difference between GPT-6 Sol, Luna, and Astra?+
Is GPT-6 Luna good enough for agentic work?+
How does GPT-6 Sol compare to Claude models?+
Is prompt caching cheaper on the GPT-6 models?+
Pros
- Sol lands within 1.1 points of Claude Fable 5's best DeepSWE score at roughly 80% lower cost per task
- Cached input reads carry a 90% discount — $0.20 per million on Sol and $0.01 on Luna
- All three GPT-6 tiers share the same 1.05M-token context window and 128K max output
- Sol roughly halves its predecessor's factual error rate and beats Claude Opus 5 on AutomationBench at 9% of the cost per task
- Free and Go users get Luna in the desktop app, so the cheapest tier is not paywalled
Cons
- Luna is not an agentic model: AutomationBench falls to 20.7% against Sol's 33.2%, and OSWorld 2.0 to 52.7% against 64.4%
- Luna's factual error rate is 7.6%, nearly double Astra's 3.9% — it needs a verification layer for customer-facing work
- Cache writes bill at 1.25x the uncached input rate, so the caching discount depends on reuse patterns
- Benchmarks are vendor-reported; only the Intelligence Index parity claim is independently measured so far
- Astra stays at $10/$50, so the 50% cut applies to the mid and low tiers, not the flagship
Keep reading