← Pooja Verma
BenchmarkedIssue · 24 Sep 2026
Model launch · about a 6-minute read

Claude Opus 5.5: a narrow lead at a lower price

Independent leaderboards put Anthropic’s new mid-priced model first, but the margin shrinks outside Anthropic’s own tests, and at max effort the token bill climbs fast.

Claude Opus 5.5 title card from Anthropic’s launch page, dated September 22, 2026

The short version

  • #1 where tested: Artificial Analysis scores it 58, five points ahead of GPT-6 Astra and Claude Fable 5.1 (53 each); Vals AI ranks it #1 of 63. independent
  • Fable-class pitch, lower price: $4 / $20 per million tokens, 40% of the list price of Fable 5.1 and GPT-6 Astra. vendor
  • Narrower than advertised: Anthropic’s Terminal-Bench 4.0 lead over Astra becomes a tie (59.6% each) in Artificial Analysis’s run. independent
  • Hungry at max effort: about 119k output tokens per task, versus about 27k for Astra. press
  • Not the leader everywhere: even Anthropic’s own system card has Astra ahead on science terminal work, workflow automation and ultra-long coding tasks. vendor

01What Anthropic announced

Opus 5.5 launched on 22 September 2026 as the first model in Anthropic’s Claude 5.5 family. Anthropic now recommends it as the starting point for most workloads, keeping Fable 5.1 for the most demanding reasoning and long-horizon agent work. vendor

02Anthropic’s scorecard

Anthropic’s launch table compares Opus 5.5 with Fable 5.1, Opus 5, GPT-6 Astra and GPT-5.6 Sol on nine benchmarks. Opus 5.5 wins seven, though on two of those (OSWorld 2.0 and Chartography) it faces only other Claude models. Astra wins the remaining two. No Google, xAI, Meta or open-weight model appears. vendor

Anthropic’s benchmark table with the Terminal-Bench 4.0 row highlighted
Anthropic’s launch table, Terminal-Bench 4.0 row highlighted: 66.4% for Opus 5.5 vs 57.9% for GPT-6 Astra, run at xhigh effort. Rival scores are the vendors’ own figures. Source: Anthropic.

The system card’s biggest coding number is 89.9% on SWE-bench Pro, against 81.2% for Fable 5.1. Nobody had reproduced it independently by 24 September. vendor

03What independent testers found

Artificial Analysis (AA), Vals AI and ARC Prize all put Opus 5.5 at or near the top. AA calls its 58 the highest index score it has recorded. independent

TestResultKind
AA Intelligence Index58, #1 of 211 (Astra 53, Fable 5.1 53)independent
AA: Humanity’s Last Exam61.4%, a record (Fable 5.1 59.1%)independent
AA: GDPval-AA v2.1 (knowledge work)1,846 Elo (Fable 5.1 1,735)independent
AA: Terminal-Bench 4.059.6%, tied with Astraindependent
Vals Index#1 of 63 at 69.69%independent
Vals: Terminal-Bench 4.0#1 of 34 at 61.62%independent
ARC-AGI-2 (verified)93.3% at high effort, 91.7% at maxindependent

Terminal-Bench shows why it matters who ran a test. Anthropic reports 66.4%, with safeguards on and a fallback model answering the 2.5% of requests they flagged. vendor Vals measured 61.62% and AA 59.6%, which is level with Astra. independent

04Coding and office work in practice

Anthropic picked its partner testimonials, so treat them as selected examples. Stripe reports one session that directed 40 stacked pull requests, all passing CI. Clio describes an 18-hour task completed unattended. Deloitte says the model caught 72% of known bugs at its lowest effort setting, against 56% for Opus 5 at high. vendor

Stripe engineer’s quote about 40 stacked pull requests passing CI
Partner quote from Cristian Rivera, Staff Software Engineer at Stripe. Source: Anthropic.

Outside reviewers largely agree. Every’s team says it is pulling staff who had switched to OpenAI’s Codex back to Claude. Simon Willison now uses it as his default in Claude Code. community

CodeRabbit’s code-review test was more mixed: 51 of 80 known bugs caught versus 49 for its production baseline, with similar precision and 49% more tokens. independent

GDPval-AA v2.1 Elo plotted against estimated cost per task for five models
GDPval-AA v2.1 (office work from 44 occupations): Elo against cost per task, one dot per effort level. Anthropic’s chart; Artificial Analysis confirms the 1,846. Source: Anthropic.

05Writing that gets to the point

Anthropic says it fixed the dense prose critics called “Claudish”: answers now lead with the most important information. vendor

Side-by-side explanations of the same billing bug by Claude Opus 5 and Claude Opus 5.5
The same bug explained by Opus 5 (left) and Opus 5.5 (right). Example chosen by Anthropic. Source: Anthropic.

Every rated its prose the most readable it has tested (Flesch-Kincaid grade 6.95). Ethan Mollick (@emollick) said it feels Fable-class but hasn’t fully solved the dense-language problem. community

06Where it falls short

Lower half of Anthropic’s table with GPT-6 Astra’s Terminal-Bench-Science win highlighted
The lower half of Anthropic’s table. GPT-6 Astra (second column from right) wins AutomationBench and Terminal-Bench-Science. Source: Anthropic.

07Cheaper per token, not always per task

On AA’s index at max effort, Opus 5.5 produced about 119k output tokens per task versus about 27k for Astra, according to OfficeChai. press AA labels it very verbose; Vals measured $22.30 per test and 72 minutes of latency. independent

Terminal-Bench 4.0 score plotted against cost per attempt for five models at each effort level
Terminal-Bench 4.0 score against cost per attempt, one dot per effort level. Anthropic’s chart, so vendor numbers. Source: Anthropic.

Willison’s trivial drawing prompt at max effort hit the 128,000-token output ceiling, at $2.56 per attempt. community One tester cited by The Neuron preferred Opus on 7 of 8 tasks, but GPT-6 Sol, at half the per-token price, finished the set in 5h51m for about $74, against 8h40m and about $213 for Opus. press

Effort is the lever. On Anthropic’s CursorBench figures, Opus 5.5 at high effort costs about $4 per task, roughly a quarter of Fable 5.1’s $17.28 at max, and still outscores it. vendor

08Safety: better behaved, more test-aware

Alignment section of Anthropic’s launch post
Anthropic’s alignment summary, which says the model “often suspects it is being evaluated.” Source: Anthropic.

Anthropic calls Opus 5.5 the strongest model on its automated behavioral audit, and says it tried to get around boundaries about 85% less often than Opus 5 or Mythos 5.1. Most cybersecurity tasks are routed to Opus 4.8. vendor

The system card also lists regressions: it is more likely to follow malicious instructions hidden in pasted text, and probes flagged evaluation awareness in 36% of audit transcripts versus 0.4% of real Claude Code use. vendor Zvi Mowshowitz argues that awareness deserves more scrutiny. community

09How it stacks up

Opus 5.5Fable 5.1GPT-6 AstraGPT-6 Sol
API price (in / out per 1M)$4 / $20$10 / $50$10 / $50$2 / $10
AA index585353—
Terminal-Bench 4.0 (Anthropic)66.4%55.8%57.9%—
Terminal-Bench 4.0 (AA)59.6%—59.6%—
GDPval-AA (Elo)1,8461,7351,542—

10Handle with care

  • All data is from the first two days. LMArena, METR, Epoch, SWE-bench and Scale SEAL had not published independent results.
  • “1,846 Elo” is a GDPval-AA score from Artificial Analysis, not an LMArena rating.
  • “40% cheaper than Opus 5” is Anthropic’s estimate for typical workloads; the list price is 20% lower.
  • AA’s 58 was measured with Anthropic’s fallback on, so some cyber and biology requests were answered by other models.
  • Partner results are testimonials selected by Anthropic, not independent studies.

11Sources

Benchmarked · the AI news, fact-checked · Facts as of 24 Sep 2026 · labels show who reported each number: vendor, independent, press, community