Claude Sonnet 5.5 and Haiku 5.5: The Complete 5.5 Family and Its Pricing Ladder (2026)
Quick verdict
Anthropic completed the Claude 5.5 family: Sonnet 5.5 shipped September 28 at $2/$10 (cache reads halved to $0.10 on October 7, cutting agentic costs ~20%), and Haiku 5.5 shipped October 7 from $0.10/$0.50 with a 1M-token context window, the first effort dial ever shipped in a Haiku-class model, and the deepest cache discount Anthropic has published at $0.01 per million reads. The family now spans a 100x price range for the same context window.
What Finished Shipping
Anthropic closed out its 5.5 generation in two steps. Claude Sonnet 5.5 arrived September 28 at $2/$10 — the family's speed-and-intelligence workhorse. Then on October 7, two things happened on the same day: Claude Haiku 5.5 launched, and Sonnet 5.5's cache reads were cut in half.
The second of those had no launch post and will move more money this quarter. Cache reads on Sonnet 5.5 fell from $0.20 to $0.10 per million tokens — now 5% of the model's input price rather than 10%. Because cached context (system prompts, repository instructions, tool definitions, conversation history) dominates token consumption in agent loops, Anthropic estimates the change makes Sonnet 5.5 roughly 20% cheaper on most agentic work.
The Ladder
The complete family, all sharing a 1M-token context window and 128K max output:
| Model | Input / Output (per 1M) | Cache reads | Position |
|---|---|---|---|
| Haiku 5.5 | $0.10 / $0.50 (≤100K) | $0.01 | High-volume, latency-sensitive work |
| Sonnet 5.5 | $2 / $10 | $0.10 (halved Oct 7) | Everyday tasks, bug fixes, docs and sheets |
| Opus 5.5 | $4 / $20 | $0.20 | Anthropic's recommended starting point |
| Fable 5.1 | $10 / $50 | $0.25 | Demanding reasoning, long-horizon agents |
| Mythos 5.1 | — | — | Fable 5.1 with permissive safeguards, trusted access only |
That is a 100x spread between the cheapest and most expensive tier — for the same context window. Hold onto that fact; it is the architectural point of the whole lineup.
Sonnet 5.5: The Workhorse, With a Cost Caveat
Sonnet 5.5 is billed as Anthropic's best speed-intelligence combination, strongest at well-scoped everyday work: bug fixes, documents, slides, spreadsheets. Default effort is high on the API and medium in Claude Code and the apps, it has no fast mode, and its minimum cacheable prompt dropped to 512 tokens from Sonnet 5's 1,024.
The independent numbers are more interesting than the marketing. On Arena's Agent Arena it ranks third with a +12.5% net improvement — but at a median $2.74 per task, roughly 73% more expensive per task than Opus 5.5's $1.58. And on Artificial Analysis's Coding Agent Index, Sonnet 5.5 at max effort tops the chart at 68 — with cost per task reaching $14.19 in Claude Code.
So: Sonnet 5.5 is the most capable mid-tier model for agentic coding, and it is not the cheap way to do that work. The cache cut is what makes it economically sensible for long sessions; without it, the token bill for the top score is steep.
Haiku 5.5: The Actually Interesting One
Every "cheapest model" launch gets the same writeup. This one has three details that change how agent systems get built.
1. It has an effort dial. Haiku 5.5 is the first Haiku-class model with adjustable effort — Low, Medium, High, Xhigh, Max, defaulting to Medium — with adaptive thinking on by default. Reasoning spend is now a knob at the cheap tier, not a fixed property. A browser-use agent can run at low effort for navigation and the same model at high effort for a policy-document classification pass.
2. It has the flagship's context window. 1M tokens, up from 200K on Haiku 4.5. A routing agent or classifier on Haiku 5.5 now carries the same context fidelity as an Opus 5.5 reasoning layer. That is not a spec-sheet detail; it changes how pipelines delegate tasks, because passing long context down a tier no longer means truncating it.
3. The price is a tenth, with a cliff. Against Haiku 4.5's flat $1.00/$5.00: $0.10 / $0.50 for prompts up to 100,000 tokens — a tenth of the price, five times the context, and cache reads at $0.01 per million, the deepest cache discount Anthropic has published on any model. Above 100,000 tokens, every line quintuples to $0.50 / $2.50. Anthropic has not said whether cached tokens count toward that threshold — budget as if they do.
Anthropic's own estimate: average workload cost about 75% below Haiku 4.5, accounting for request sizes and tokenizer changes. It is also the fastest model Anthropic has shipped at standard speed.
Where the benchmarks stand. Haiku 5.5 puts up 72.4% on OSWorld 2.1 and 39.2% on Terminal-Bench 4.0 at max effort — vendor-reported figures from the system card, with independent evaluation still pending. Its independent Intelligence Index score is 43, up 26 points year over year. On its Code Arena debut it scored 1587, landing 30th. Anthropic's own framing is the right one to keep: Sonnet 5.5 and Opus 5.5 remain better choices for complex agentic coding; Haiku 5.5 is for the narrow, high-volume work "that might otherwise have been cost-prohibitive" — compaction, summarization, classification, extraction, routing, database queries, and subagent tasks under a Sonnet or Opus orchestrator.
On safety, the tier moved with the family: cyber safeguards are stricter than Haiku 4.5's (defensive work allowed, penetration testing still blocked), and biology safeguards match those on Sonnet 5, Sonnet 5.5 and Opus 5. Alignment evaluations show far fewer misaligned behaviors than Haiku 4.5.
The Architectural Point
Put the three facts together and a design pattern becomes obvious. Anthropic now sells one context window at five prices, and the cheapest tier has both an effort dial and a 1M window. That means the subagent pattern — cheap model does the volume work, expensive model does the reasoning — can now run with full context fidelity all the way down. Compaction and summarization, the two jobs whose entire purpose is dealing with long context, are now viable at $0.10 input and $0.01 cache reads.
The same logic that made GPT-6 Luna a worker rather than a planner applies here: the win is not that the cheap model is smart, it is that routing work to it no longer costs you context or configurability.
Migration Is Not Drop-In
Anthropic's release notes are explicit, and teams should read them before flipping production traffic:
- Adaptive thinking is on by default — responses may begin with thinking blocks even if your code never configured thinking. Any parser assuming the first block is text will break
- The tokenizer changed — the same prompt text yields different token counts, invalidating cost models built on fixed estimates
- Requests using
budget_tokensneed reworking, and tool-use loops that serialize or replay thinking blocks need checking - Code written for Haiku 4.5 may not function correctly against Haiku 5.5 despite the model ID being the only obvious change
Also note the deprecation clock: Claude Sonnet 4.5 was deprecated September 30 and retires November 30, 2026.
Who Should Use Which
| Workload | Tier |
|---|---|
| Classification, extraction, routing, summarization, compaction | Haiku 5.5 — the price gap to Sonnet is 20x |
| Subagents under a Sonnet/Opus orchestrator | Haiku 5.5 — same 1M context, effort dial for quality control |
| Long-running agent sessions where context is reused | Sonnet 5.5 — the halved cache read is the whole argument |
| Everyday coding, documents, slides, sheets | Sonnet 5.5 |
| Complex agentic coding where the best score matters more than cost | Opus 5.5 — cheaper per task than Sonnet 5.5 at max effort, and it leads the Intelligence Index |
| Hardest reasoning, long-horizon autonomy | Fable 5.1 |
Summary
The 5.5 family is complete, and the interesting story is not any single model — it is that Anthropic now sells the same 1M-token context window across a 100x price range, and gave the cheapest tier a thinking dial. That combination is what makes aggressive subagent architectures viable without losing context fidelity.
The caveats are the familiar ones from this era of rapid iteration: vendor-reported agentic scores awaiting independent evaluation, a pricing cliff at 100,000 tokens, and a migration that breaks naive parsers. The honest summary: Sonnet 5.5 is the workhorse whose economics improved the most this week, and Haiku 5.5 is the tier that will quietly absorb the most tokens.
Related Articles
Frequently Asked Questions
What does the full Claude 5.5 lineup cost?+
What is the 100,000-token cliff in Haiku 5.5's pricing?+
How much cheaper is Haiku 5.5 than Haiku 4.5?+
What is the first-Haiku effort dial?+
Do I need to change code when migrating to Haiku 5.5?+
Pros
- Haiku 5.5 costs a tenth of Haiku 4.5 on the same prompt shape while offering five times the context window
- Every tier in the family now shares a 1M-token context window and 128K max output
- Sonnet 5.5's cache reads halved to $0.10 — 5% of input price, the single change most likely to move real bills
- Haiku 5.5 is the first Haiku with an effort dial (Low to Max), making reasoning spend adjustable at the cheap tier
- Haiku 5.5's cache reads of $0.01 per million are the deepest discount Anthropic has published
Cons
- Haiku 5.5's headline price quintuples above 100,000 tokens: $0.50 input and $2.50 output
- Anthropic has not said whether cached tokens count toward that 100K threshold — budget as if they do
- Sonnet 5.5 tops the Coding Agent Index but at up to $14.19 per task, well above Opus 5.5's cost efficiency
- Haiku 5.5's agentic benchmarks (72.4% OSWorld 2.1, 39.2% Terminal-Bench 4.0) are self-reported; independent evaluation is pending
- Migration is not drop-in: adaptive thinking is on by default, so responses may lead with thinking blocks and token counts shift with the new tokenizer
Keep reading