Anthropic released Claude Opus 5.5 on 22 September 2026, the first model in the 5.5 family. It is priced at $4 per million input tokens and $20 per million output tokens — 40% below Opus 5 — while scoring close to Fable 5.1. The launch landed on the same day as OpenAI’s new models, shifting the contest between labs onto price.
Key takeaways
- Pricing: $4 per million input tokens and $20 per million output, 40% below Opus 5
- Terminal-Bench 4.0: 66.4% against 55.8% for Fable 5.1 and 52.3% for Opus 5
- Output generation more than 30% faster than Opus 5
- 85% less likely than Opus 5 to attempt to circumvent set boundaries in the behavioural audit
- Cybersecurity tasks are routed to the older Opus 4.8
The lead shows up mainly in agentic work
Opus 5.5 opens the widest gap where it runs in a terminal across many steps without supervision. On Terminal-Bench 4.0 it is more than ten percentage points ahead of Fable 5.1, and on AutomationBench it scores nearly one and a half times its predecessor.
| Benchmark | Opus 5.5 | Fable 5.1 | Opus 5 |
|---|---|---|---|
| Terminal-Bench 4.0 | 66.4% | 55.8% | 52.3% |
| CursorBench 4.0 | 57.8% | 51.8% | — |
| Humanity’s Last Exam (with tools) | 67.7% | 65.6% | — |
On knowledge tasks the differences narrow — on OSWorld 2.0 roughly a point separates the two models. Anthropic illustrates practical behaviour with a code migration covering 680,000 lines, finished in under a day.
Price matters more than the leaderboard
In agentic work a single task consumes hundreds of thousands of tokens, so the per-million rate decides whether a project stays an experiment or reaches production.
Opus 5.5 pricing and the model identifier at cloud providers:
Safeguards and deliberate limits
Anthropic calls Opus 5.5 the strongest-performing model in its automated behavioural audit?Behavioural audit: An automated battery of attempts in which the model is deliberately pushed to circumvent its constraints. It measures how often it actually tries.. External evaluation was carried out by Frontier Design and METR. The model routes cybersecurity tasks to Opus 4.8, and in biology the safeguards match those used for Fable 5.1.
Why it matters
Until now, picking a frontier model came down to asking which one was strongest. Opus 5.5 moves the centre of gravity to cost per completed task — a metric that means more to teams building agents than a leaderboard position. If the pace of price cuts holds, inference margin stops being a source of advantage, and the only differentiator left is how much work a model closes without a human.
What’s next
- Claude Sonnet 5.5 and Haiku 5.5 announced for the coming weeks according to Anthropic’s statement
- Cybersecurity tasks stay routed to Opus 4.8 — unlocking them requires going through the Cyber Verification Program
- The claim of an 85% lower tendency to circumvent boundaries awaits independent verification beyond Frontier Design and METR





