Claude Opus 5.5 Is Out: Price, Benchmarks, Real Cost

Claude Opus 5.5 launch: Anthropic ships the model the leaks predicted, at $4 input and $20 output

Claude Opus 5.5 shipped on September 22, 2026, as claude-opus-5-5, at $4 per million input tokens and $20 per million output. Anthropic says it performs at the level of Claude Fable 5.1 on most tasks, and that it costs 40% less to run than Opus 5, which is a bigger drop than the 20% price cut on the label. That gap is the most interesting thing about this release, and almost nobody is reporting it.

Key takeaways

  • Price per token fell 20%. Cost per task fell 40%. Those are different numbers and Anthropic publishes both. The extra 20 points come from a lower default effort level, a cache read price cut of 60%, and lower serving compute. If you pin effort to high, you give most of it back.
  • Every headline benchmark score was run at max or xhigh effort. The model ships defaulting to medium. Anthropic says so in a footnote under its own chart. The 66.4% on Terminal-Bench 4.0 is not the number you get out of the box.
  • Cache reads are now $0.20 per million, which is 0.05x the input price. Every other Anthropic model except the Fable line charges 0.1x. On a cache-heavy agent loop this one line matters more than the headline price.
  • Anthropic's own chart shows GPT-6 Astra ahead of Opus 5.5 on two of nine benchmarks: AutomationBench (41.4% against 40.0%) and Terminal-Bench-Science (64.6% against 58.7%). The launch was not a clean sweep and Anthropic did not pretend it was.
  • Thinking can no longer be turned off. Sending thinking: {"type": "disabled"} now returns a 400. Effort is the only lever, and a saved effort setting does not carry over to a newly released model.
  • The leak was right about almost everything. The pricing table that circulated on X the day before launch, which we described as unsourced and suspiciously tidy, matched the shipped numbers exactly, cache write included. We were wrong to lean as hard as we did on that. The scorecard is below.
  • Sonnet 5.5 shipped on September 28 at $2 and $10, and Anthropic reports it at 70.6% on Terminal-Bench 4.0, above Opus 5.5's 66.4%. Haiku 5.5 is still "in the coming weeks". There is no Fable 5.5, whatever X says.

Published September 22, 2026, before the launch. Rewritten September 23 after it. Updated October 2 with Sonnet 5.5, which shipped on September 28, and Opus 5.5's first independent index results. The original version of this page was a rumor check, written while Opus 5.5 was still unannounced. It is preserved as a scorecard further down rather than quietly deleted, because how the leak performed is part of the story. Every price, model ID and benchmark figure here is read off Anthropic's own launch page, pricing table and migration docs. Anything not published by Anthropic is labelled.

What actually shipped

Claude Opus 5.5 is the first model in what Anthropic calls the Claude 5.5 family. The model ID is claude-opus-5-5, and the same string works on Amazon Bedrock, Google Cloud and Microsoft Foundry. It is live in the API, on claude.ai and in Claude Code.

In Claude Code it is already the default Opus model as of v2.1.280, and the same release moved Pro and Team Standard plans off Sonnet and onto Opus by default, which Max and Enterprise plans already were. If you have not changed anything, you are probably running it right now.

One migration detail worth knowing before it bites you: a saved /effort level does not carry across to a newly released model. Opus 5.5 starts at its own default, which is medium, until you set it again.

Claude Opus 5.5 pricing, and the rest of the lineup

The published prices, per million tokens. Cache write is quoted for the 5 minute TTL; the 1 hour TTL costs more.

Model Input Output Cache read Context
Opus 5.5 $4 $20 $0.20 (0.05x) 1M
Opus 5 $5 $25 $0.50 (0.1x) 1M
Fable 5.1 $10 $50 $0.25 (0.025x) 1M
Sonnet 5.5 $2 $10 $0.20 (0.1x) 1M
Sonnet 5 $2 $10 $0.20 (0.1x) 1M
Haiku 4.5 $1 $5 $0.10 (0.1x) 200K

Batch requests are half price on input and output across the range. There is also a Fast mode in research preview, API only, which runs Opus 5.5 at $8 in and $40 out.

The 20% that is really 40%

This is the part worth slowing down for, because two different numbers are circulating in the same sentence and they measure different things.

The list price fell 20%: input went from $5 to $4, output from $25 to $20. Straight arithmetic, no interpretation needed.

The cost of doing a job fell further. Anthropic's own wording is that "at default settings it will cost 40% less than Opus 5 on typical workloads." Three published changes close that gap, and none of them are visible on a price list.

Lever Opus 5 Opus 5.5 Effect on a real bill
List price per token $5 / $25 $4 / $20 20% off, unconditionally
Default effort level high medium Fewer thinking and output tokens per task, if you leave it alone
Cache read multiplier 0.1x input 0.05x input $0.50 to $0.20, a 60% cut on the line item agents hit most
Serving compute baseline lower Anthropic's stated reason the price could fall at all
Thinking at a fixed effort level baseline more Works against you, most at xhigh and max

That last row is the one to read twice. Anthropic states plainly that at the same effort setting, Opus 5.5 "tends to think more per turn than Claude Opus 5, most of all at xhigh and max." So the 40% is a default-settings number. Pin effort to high to match what you were running on Opus 5, and you keep the 20% list price cut and the cache saving, but you hand back the effort-default saving and pay for extra thinking on top.

The practical version: if you migrate by changing the model string and nothing else, you probably see something near 40%. If you migrate by also forcing effort up to match your old benchmark numbers, you might see very little. Measure your own workload before you promise finance a number.

Benchmarks, and the footnote that changes all of them

Anthropic published nine comparisons on the launch page. Here they are, with the competitor columns Anthropic itself chose to show.

Benchmark Opus 5.5 Fable 5.1 Opus 5 GPT-6 Astra
Terminal-Bench 4.0 66.4% 55.8% 52.3% 57.9%
FrontierCode v1.1 54.4% 50.3% 48.0% 53.3%
CursorBench 4.0 57.8% 51.8% 46.6% not shown
GDPval-AA v2.1 (Elo) 1846 1735 1708 1542
OSWorld 2.0 (computer use) 81.8% 80.7% 74.0% not shown
Humanity's Last Exam (tools) 67.7% 65.6% 63.6% 57.2%
Chartography (tools) 89.0% 88.4% 83.4% not shown
AutomationBench 40.0% 31.4% 26.9% 41.4%
Terminal-Bench-Science 0.1 58.7% 52.6% 29.0% 64.6%

Now the footnote, quoted from under Anthropic's own chart: "Unless otherwise noted, all Claude Opus 5.5 results use adaptive thinking at max effort. Terminal-Bench 4.0 results are reported for Claude Opus 5.5 at xhigh effort and GPT-6 Astra at high effort."

Read those two sentences together

The benchmark numbers are produced at max or xhigh effort. The model ships defaulting to medium. And at a fixed effort level it thinks more per turn than Opus 5 did.

So the 66.4% and the 40% saving are two different configurations of the same model. You can have the benchmark behaviour or you can have the default-settings bill. Publishing both on one page is fair; quoting them in one sentence, as most of the coverage is doing, is not.

Two more honest notes on that table. Anthropic did not publish GPQA, AIME, MMMLU, tau-bench or ARC-AGI numbers for Opus 5.5 at all, so any table you see carrying those under an Opus 5.5 label has borrowed them from an older model. And Anthropic's own chart shows GPT-6 Astra winning two of the nine, which is more candour than most launch pages manage.

What changed in how it behaves

The score table is the least interesting part of this release. The behavioural notes in the migration docs matter more if you actually run this thing in production.

  • Thinking cannot be disabled. thinking: {"type": "disabled"} returns a 400. Effort is the only control left, and it now spans medium as the default up through xhigh and max.
  • Communication style was deliberately changed. Anthropic says it "puts the most important information up front", is "less likely to use jargon or idiosyncratic phrases", and "follows the writing rules you give it". That last clause is the one worth testing if you have ever fought a model over a style guide.
  • Computer use needs the new toolset. On the API and Google Cloud you must move to computer_toolset_20260801; the older computer_20251124 is rejected. Bedrock still accepts the old one for now.
  • Vision got specifically better at charts. Anthropic claims it "reads values off dense charts and layout-dependent visuals much more precisely without tools", which is a narrow claim and an easy one to verify yourself.
  • Long-horizon plumbing. Task budgets, mid-conversation system messages, compact-on-demand in beta, and thinking blocks preserved across a conversation.

On the claim going around that Opus 5.5 has stopped writing em dashes: that is a user observation, not an Anthropic statement. We looked through the launch page, the migration docs and the model card and found nothing about punctuation. It is plausible as a side effect of "follows the writing rules you give it", and it is not something Anthropic has said.

Usage limits on Pro, Max and Team

Anthropic's own sentence is that it is "increasing five-hour usage limits on Pro, Max, Team, and seat-based Enterprise plans", plus a one-time rate limit reset that subscribers can save and spend whenever they like.

The specific figures everyone is quoting, a 20% higher five-hour limit and Opus 5.5 going "about 25% further" inside it, are consistent across the coverage but we could not pin either one to a verbatim Anthropic sentence. Treat them as reported rather than published. The mechanism is not mysterious: a higher quota multiplied by a cheaper per-task cost compounds.

The rumor scorecard: what the leak got right

The day before launch, this page argued that the circulating claims were thin. Here is how they actually landed. We are keeping our own wrong calls in the table rather than editing them out.

Claimed on September 21 What shipped Verdict
Model ID claude-opus-5-5 claude-opus-5-5 Right
Named Opus 5.5, not 5.1 or 5.2 Opus 5.5, first of a 5.5 family Right, and we warned it would be wrong
$4 input, $20 output $4 input, $20 output Right
Cache read $0.20 $0.20 Right
Cache write $5 $5 at the 5 minute TTL Right, and we refused to print it
Releases September 22 September 22 Right
Performs at Fable 5.1 level Anthropic's own phrasing, on most tasks Right
Context 900K to 1M 1M, unchanged from Opus 5 Right, and still not an upgrade
Codename claude-wafer-eap Never appeared in any public artifact, before or after Unresolved

What we got wrong

Eight of nine. The one line we leaned on hardest was that the price list looked like a fabrication because every figure was a clean 20% off a published Opus 5 number. That reasoning was sound and the conclusion was wrong: the figures were tidy because the actual price cut was tidy.

The deeper mistake was a category error. We were grading the rumor on the strength of its evidence, which was genuinely weak, and then letting that slide into a verdict on whether it was true. Those are different questions. A claim with no visible artifact behind it can still be an accurate claim from someone who saw the real thing, and in this case it plainly was.

What we would keep: the refusal to print the cache write figure we could not check, the correction of the "88% on Polymarket" number that was actually 81.8%, and the insistence on separating what Anthropic had published from what an anonymous account asserted. The method held up. The prediction did not.

One thing did go the way the pattern suggested. The Honeycomb leak we covered in July reached a real artifact, a model string inside Cursor's production UI, and it also preceded a real launch. Codename chatter with no artifact behind it has a worse record, which is why the scorecard above is worth keeping rather than quietly overwriting.

Should you switch?

For agentic coding and computer use, the case is straightforward: it is the top score on six of the nine benchmarks Anthropic published, and it is cheaper per task than the model it replaces. In Claude Code the decision has already been made for you.

The independent results that have landed since launch back that up. Opus 5.5 is first on the Epoch Capabilities Index at 167.35, a lead of 0.84 over GPT-6 Astra that is too small to call a win on its own. On the Artificial Analysis Intelligence Index, Anthropic holds the top six slots, with Opus 5.5 at 58 at max effort. The one caveat is the arena boards: Opus 5.5 does not lead Arena's text or WebDev rankings, whatever the trending posts say.

The bigger change is underneath it. Claude Sonnet 5.5 shipped on September 28 at half the price, $2 and $10, with the same 1M context and 128K output. Anthropic reports it at 70.6% on Terminal-Bench 4.0, above Opus 5.5's 66.4%, and says it writes output more than 30% faster than Sonnet 5. It sits at 56 on the Artificial Analysis index and third on Epoch at 165.2, just above Fable 5.1. Before you move a terminal-heavy coding workload to Opus 5.5, run it on Sonnet 5.5 too.

The migration is a model string, with three caveats that will cost you an afternoon if you skip them. Move computer use to the new toolset. Remove any code that disables thinking, because it now errors. And re-set your effort level explicitly rather than assuming your old one carried over, then measure the bill before and after on your own traffic instead of trusting either the 20% or the 40%.

If you are choosing between vendors rather than upgrading within one, the comparison changed twice this week: OpenAI shipped GPT-6 Sol and Luna within a day of this launch and GPT-6.1 Sol a week later at the same $2 and $10, and if you want Sol driving your terminal rather than Claude, we have a walkthrough of running an OpenAI model inside Claude Code. We keep a running comparison of the frontier models with the current pricing and the independent benchmark picture, which is the better page for that question. For how the previous generation behaved in real work rather than on a chart, our Fable 5 review and the account of what Opus 5 did inside a live security research chain are more useful than any benchmark table.

Find out what the switch actually costs you

A 20% list price cut and a 40% cost-per-task claim are both true and neither one is your number. Yours depends on your cache hit rate, your effort setting, how much of your spend is output tokens, and whether your prompts were tuned to a model that no longer exists in the same configuration.

Valletta Software Development builds the layer that answers this in a day rather than a quarter: a provider-agnostic routing layer, an eval suite that runs in CI against your own tasks, per-route cost and latency telemetry, and a migration path that does not touch product code. Tell us what you are running and what you are spending, and we will tell you what Opus 5.5 changes for you specifically.

Book a scoping call

FAQ

Is Claude Opus 5.5 out?

Yes. Anthropic released Claude Opus 5.5 on September 22, 2026. The model ID is claude-opus-5-5 and it is available through the Claude API, claude.ai, Claude Code, Amazon Bedrock, Google Cloud and Microsoft Foundry. It is now the default Opus model in Claude Code.

How much does Claude Opus 5.5 cost?

$4 per million input tokens and $20 per million output tokens, with cache reads at $0.20 per million and cache writes from $5. That is 20% below Opus 5 on list price. Anthropic separately says that at default settings a typical workload costs 40% less than on Opus 5, because the default effort level dropped from high to medium and cache reads were cut by 60%.

Why is Opus 5.5 40% cheaper if the price only dropped 20%?

Because the two figures measure different things. The 20% is the list price per token. The 40% is the cost of finishing a typical task at default settings, which also benefits from the lower default effort level, the cache read multiplier dropping from 0.1x to 0.05x of input, and lower serving compute. If you override effort back up to high or max, you give much of that back, since at a fixed effort level Opus 5.5 thinks more per turn than Opus 5 did.

Is Claude Opus 5.5 better than Fable 5.1?

On the nine benchmarks Anthropic published, Opus 5.5 beats Fable 5.1 on all of them, and Anthropic describes it as performing at Fable 5.1's level on most tasks while costing 60% less per token. Those benchmark numbers are run at max or xhigh effort, not the medium default, so treat them as a ceiling rather than an everyday expectation.

Does Opus 5.5 beat GPT-6 Astra?

Not everywhere, and Anthropic's own chart says so. Opus 5.5 leads on Terminal-Bench 4.0 (66.4% against 57.9%), FrontierCode (54.4% against 53.3%), GDPval-AA and Humanity's Last Exam. GPT-6 Astra leads on AutomationBench (41.4% against 40.0%) and Terminal-Bench-Science (64.6% against 58.7%). OpenAI also shipped GPT-6 Sol and Luna a day later, which those numbers do not cover.

What is the Claude Opus 5.5 context window?

1 million tokens, with 128,000 tokens of maximum output, or 300,000 through the Batch API with a beta header. Both are unchanged from Opus 5, Sonnet 5 and Fable 5.1, and Sonnet 5.5 matches them. The pre-launch rumor of a 900K to 1M context window was accurate and was never an upgrade.

Did Opus 5.5 stop using em dashes?

Users report that it does, and Anthropic has not said so. There is no mention of dashes or punctuation in the launch page, the migration documentation or the model card. Anthropic does say the model "follows the writing rules you give it" and avoids "idiosyncratic phrases", which would explain the observation without confirming it.

Is Claude Sonnet 5.5 out, and when is Haiku 5.5 coming?

Sonnet 5.5 is out: it shipped on September 28, 2026, as claude-sonnet-5-5, at $2 per million input tokens and $10 per million output. Haiku 5.5 is still coming. Anthropic says it "will join the Claude 5.5 family in the coming weeks" and has not given a date or a price.

What broke in the migration from Opus 5?

Three things. Thinking can no longer be disabled, so a request carrying thinking: {"type": "disabled"} now returns a 400. Computer use on the API and Google Cloud requires the computer_toolset_20260801 toolset, and the older computer_20251124 is rejected. And a saved effort setting does not carry over to a newly released model, so Opus 5.5 starts at its own medium default until you set it again.

Is Claude Fable 5.5 out?

No. As of October 2, 2026, the newest Fable model is Fable 5.1, released September 1, and Anthropic has announced no Fable 5.5. Posts on X crediting outputs to "Fable 5.5" are mislabelled. The Claude 5.5 family so far is Opus 5.5 and Sonnet 5.5, with Haiku 5.5 to follow.

Vibe coded an app that needs to become a product?

We audit AI generated codebases and bring them to production quality. Book a free 30 minute call and get a senior review of your project.

Valletta.Software - Top-Rated Agency on 50Pros

Talk to the engineers behind this blog

Valletta Software builds and staffs dedicated development teams for companies across the EU and US. Tell us about your project and get a reply within one business day.