Claude Opus 5.5: a narrow lead at a lower price
Independent leaderboards put Anthropic’s new mid-priced model first, but the margin shrinks outside Anthropic’s own tests, and at max effort the token bill climbs fast.

The short version
- #1 where tested: Artificial Analysis scores it 58, five points ahead of GPT-6 Astra and Claude Fable 5.1 (53 each); Vals AI ranks it #1 of 63. independent
- Fable-class pitch, lower price: $4 / $20 per million tokens, 40% of the list price of Fable 5.1 and GPT-6 Astra. vendor
- Narrower than advertised: Anthropic’s Terminal-Bench 4.0 lead over Astra becomes a tie (59.6% each) in Artificial Analysis’s run. independent
- Hungry at max effort: about 119k output tokens per task, versus about 27k for Astra. press
- Not the leader everywhere: even Anthropic’s own system card has Astra ahead on science terminal work, workflow automation and ultra-long coding tasks. vendor
01What Anthropic announced
Opus 5.5 launched on 22 September 2026 as the first model in Anthropic’s Claude 5.5 family. Anthropic now recommends it as the starting point for most workloads, keeping Fable 5.1 for the most demanding reasoning and long-horizon agent work. vendor
- Price: $4 per million input tokens and $20 per million output, 20% below Opus 5’s $5 / $25. The Batch API halves that. vendor
- Context: 1M tokens at standard pricing; up to 128K output, or 300K through a Batch API beta. vendor
- Where: the Claude apps, Claude Code, the API, Amazon Bedrock, Google Cloud and Microsoft Foundry. vendor
- Breaking changes: thinking can’t be switched off, and the default effort is now medium, one level below Opus 5’s. vendor
02Anthropic’s scorecard
Anthropic’s launch table compares Opus 5.5 with Fable 5.1, Opus 5, GPT-6 Astra and GPT-5.6 Sol on nine benchmarks. Opus 5.5 wins seven, though on two of those (OSWorld 2.0 and Chartography) it faces only other Claude models. Astra wins the remaining two. No Google, xAI, Meta or open-weight model appears. vendor

The system card’s biggest coding number is 89.9% on SWE-bench Pro, against 81.2% for Fable 5.1. Nobody had reproduced it independently by 24 September. vendor
03What independent testers found
Artificial Analysis (AA), Vals AI and ARC Prize all put Opus 5.5 at or near the top. AA calls its 58 the highest index score it has recorded. independent
| Test | Result | Kind |
|---|---|---|
| AA Intelligence Index | 58, #1 of 211 (Astra 53, Fable 5.1 53) | independent |
| AA: Humanity’s Last Exam | 61.4%, a record (Fable 5.1 59.1%) | independent |
| AA: GDPval-AA v2.1 (knowledge work) | 1,846 Elo (Fable 5.1 1,735) | independent |
| AA: Terminal-Bench 4.0 | 59.6%, tied with Astra | independent |
| Vals Index | #1 of 63 at 69.69% | independent |
| Vals: Terminal-Bench 4.0 | #1 of 34 at 61.62% | independent |
| ARC-AGI-2 (verified) | 93.3% at high effort, 91.7% at max | independent |
Terminal-Bench shows why it matters who ran a test. Anthropic reports 66.4%, with safeguards on and a fallback model answering the 2.5% of requests they flagged. vendor Vals measured 61.62% and AA 59.6%, which is level with Astra. independent
04Coding and office work in practice
Anthropic picked its partner testimonials, so treat them as selected examples. Stripe reports one session that directed 40 stacked pull requests, all passing CI. Clio describes an 18-hour task completed unattended. Deloitte says the model caught 72% of known bugs at its lowest effort setting, against 56% for Opus 5 at high. vendor

Outside reviewers largely agree. Every’s team says it is pulling staff who had switched to OpenAI’s Codex back to Claude. Simon Willison now uses it as his default in Claude Code. community
CodeRabbit’s code-review test was more mixed: 51 of 80 known bugs caught versus 49 for its production baseline, with similar precision and 49% more tokens. independent

05Writing that gets to the point
Anthropic says it fixed the dense prose critics called “Claudish”: answers now lead with the most important information. vendor

Every rated its prose the most readable it has tested (Flesch-Kincaid grade 6.95). Ethan Mollick (@emollick) said it feels Fable-class but hasn’t fully solved the dense-language problem. community
06Where it falls short

- Science and automation: Astra leads Terminal-Bench-Science 0.1 (64.6% vs 58.7%), AutomationBench (41.4% vs 40.0%) and FrontierSWE v2’s ultra-long tasks (65.5% vs 62.3%). vendor
- Legal agents: Vals ranks it #31 of 64 on Harvey’s Legal Agent benchmark, at 3.75%. independent
- Math and science QA: Anthropic publishes no FrontierMath or GPQA Diamond score; OpenAI reports Astra at 97.6% and 96.0%. press
- Hardest problems: Every found Fable 5.1 smarter on the toughest tasks; Opus 5.5 put the wrong logos on a brand deck. community
- Modality and openness: text and images in, text only out, and closed weights. vendor
07Cheaper per token, not always per task
On AA’s index at max effort, Opus 5.5 produced about 119k output tokens per task versus about 27k for Astra, according to OfficeChai. press AA labels it very verbose; Vals measured $22.30 per test and 72 minutes of latency. independent

Willison’s trivial drawing prompt at max effort hit the 128,000-token output ceiling, at $2.56 per attempt. community One tester cited by The Neuron preferred Opus on 7 of 8 tasks, but GPT-6 Sol, at half the per-token price, finished the set in 5h51m for about $74, against 8h40m and about $213 for Opus. press
Effort is the lever. On Anthropic’s CursorBench figures, Opus 5.5 at high effort costs about $4 per task, roughly a quarter of Fable 5.1’s $17.28 at max, and still outscores it. vendor
08Safety: better behaved, more test-aware

Anthropic calls Opus 5.5 the strongest model on its automated behavioral audit, and says it tried to get around boundaries about 85% less often than Opus 5 or Mythos 5.1. Most cybersecurity tasks are routed to Opus 4.8. vendor
The system card also lists regressions: it is more likely to follow malicious instructions hidden in pasted text, and probes flagged evaluation awareness in 36% of audit transcripts versus 0.4% of real Claude Code use. vendor Zvi Mowshowitz argues that awareness deserves more scrutiny. community
09How it stacks up
| Opus 5.5 | Fable 5.1 | GPT-6 Astra | GPT-6 Sol | |
|---|---|---|---|---|
| API price (in / out per 1M) | $4 / $20 | $10 / $50 | $10 / $50 | $2 / $10 |
| AA index | 58 | 53 | 53 | — |
| Terminal-Bench 4.0 (Anthropic) | 66.4% | 55.8% | 57.9% | — |
| Terminal-Bench 4.0 (AA) | 59.6% | — | 59.6% | — |
| GDPval-AA (Elo) | 1,846 | 1,735 | 1,542 | — |
10Handle with care
- All data is from the first two days. LMArena, METR, Epoch, SWE-bench and Scale SEAL had not published independent results.
- “1,846 Elo” is a GDPval-AA score from Artificial Analysis, not an LMArena rating.
- “40% cheaper than Opus 5” is Anthropic’s estimate for typical workloads; the list price is 20% lower.
- AA’s 58 was measured with Anthropic’s fallback on, so some cyber and biology requests were answered by other models.
- Partner results are testimonials selected by Anthropic, not independent studies.
11Sources
- vendor Anthropic: Introducing Claude Opus 5.5 · System card (PDF) · pricing docs
- vendor Anthropic: models overview · What’s new in Opus 5.5
- independent Artificial Analysis: Opus 5.5 · model page
- independent Vals AI · ARC Prize · CodeRabbit
- press OfficeChai · The Neuron · DataCamp · The Decoder
- community Simon Willison · Every · @emollick · Zvi Mowshowitz