Claude Sonnet 5.5: near-Opus scores at half the token price
Independent tests confirm it sits two points behind Opus 5.5. The catch: at max effort it writes more output tokens per task than any model Artificial Analysis has measured.

The short version
- Near-Opus, confirmed: Artificial Analysis scores it 56 on its Intelligence Index, two points behind Opus 5.5’s 58 and #2 at launch. independent
- Same price as Sonnet 5: $2 / $10 per million input/output tokens, half of Opus 5.5’s $4 / $20 per token. vendor
- Token-hungry at max: about 193K output tokens per task, the most AA has measured, so a max-effort task costs about 50% more than with Sonnet 5. independent
- Cheaper at lower effort: at Medium it beats Sonnet 5’s best Terminal-Bench score for less than a tenth of the cost per task, per Anthropic. vendor
- Opus still leads on hard work: Anthropic itself says Opus 5.5 “remains clearly stronger” at complex, open-ended tasks. vendor
01What Anthropic announced
Sonnet 5.5 launched on 28 September as the second model in the Claude 5.5 family, six days after Opus 5.5. Anthropic pitches it as a faster, cheaper complement to Opus, “strongest at well-scoped everyday tasks, fixing bugs, and creating polished documents, slides, and spreadsheets.” vendor
- Specs: 30%+ faster output than Sonnet 5, a 1M-token context window, a June 2026 knowledge cutoff. vendor
- Effort: five levels, low to max. Default is Medium in the Claude apps and Claude Code, High on the API. vendor
- Safety: the first Sonnet with Opus-style cyber safeguards and anti-distillation classifiers. vendor
02Anthropic’s scorecard
On Anthropic’s table, Sonnet 5.5 sits within a few points of Opus 5.5 almost everywhere and beats it on one row: Terminal-Bench 4.0 (multi-step command-line work), at 70.6% vs 66.4% for Opus at Xhigh effort. Sonnet 5 scored 10.3% on the same test. vendor

Elsewhere Opus keeps a small lead: CursorBench 4.0 (55.5% vs 57.8%), OSWorld 2.1 (80.1% vs 81.8%) and FrontierCode 1.1 (52.1% at Xhigh vs 54.4%). vendor
Sonnet 5’s Terminal-Bench score looks suspiciously low, but Artificial Analysis’s own run also shows a jump of about 50 points. independent
03What independent testers found
Artificial Analysis (AA) backs up the near-Opus claim: 56 on its Intelligence Index at max effort, two points behind Opus 5.5 and 18 ahead of Sonnet 5. independent

| Test | Result | Kind |
|---|---|---|
| AA Intelligence Index (max) | 56, #2 at launch (Opus 5.5 58) | independent |
| Terminal-Bench 4.0, AA’s run | 64% (Opus 5.5, GPT-6 Astra 60%) | independent |
| GDPval-AA v2.1 (Elo) | 1844 (Opus 5.5 1846) | independent |
| AA-Briefcase v1.1 (Elo) | 1811 (Opus 5.5 1822) | independent |
| AA-Omniscience accuracy | 54% (Opus 5.5 66%) | independent |
04The catch: tokens
Price per token is not price per task. At max effort, AA measured about 193K output tokens per Intelligence Index task, the highest it has seen: roughly 60% more than Opus 5.5 or Sonnet 5 at max, and about 7× GPT-6 Astra. independent

A max-effort task cost $7.60, about 50% more than with Sonnet 5. independent A chart of AA’s data from @haider1 puts Opus 5.5 (max) at 119K tokens, GPT-6 Sol at 31K and Astra at 27K. community
A widely shared claim of “150% more tokens” than Opus is wrong: AA’s figure is about 60%, and only at max effort, as @neamtuz noted. community
05Cheaper or pricier? It depends on effort
Anthropic says Sonnet 5.5 “costs up to 30% less per task” than Sonnet 5 vendor; AA says about 50% more independent. Both hold. They describe different effort settings.

| Effort | AA Intelligence Index | Cost per AA task |
|---|---|---|
| Max | 56 | $7.60 |
| Xhigh | 52 | $2.74 |
| High | 47 | $1.08 |
| Medium | 41 | $0.59 |
| Low | 36 | $0.41 |
AA calls High the most competitive setting: very narrowly behind GPT-6 Sol on intelligence at effectively the same cost per task. independent
06Real use cases and demos
In a Claude Code bug-fix race posted by Anthropic’s Boris Cherny (median of 7 runs), Sonnet 5.5 finished in 12.3 s for $0.05 and 13.3K tokens, against 17.5 s, $0.08 and 18.2K for Sonnet 5. vendor

- GitHub: “matched Sonnet 5 on coding tasks while using fewer steps, tokens, and tool calls.” independent
- Cognition: 64.4% on FrontierCode 1.1 in its Devin harness, vs 56.2% for Sonnet 5. A different setup from Anthropic’s, so not comparable. independent
- Quoted by Anthropic: Base44, apps level with Opus 5 in 3.6 iterations vs 7.7; Zendesk, tickets 20% faster; Balyasny, about 121K tokens per answer vs 497K for Sonnet 5; Box, 2.4× faster. Two experts judged a 10-slide review it drafted “ready to send as is.” vendor

On Zapier’s test it scored 44.7% at Max vs 40% for Opus 5.5, at about $1.08 a task vs $1.28. It was strongest in Operations (59%) and weakest in HR (31.7%). independent
Builders posted one-shot games and sims, including a Fall Guys clone and a Starship sim from early tester @MatthewBerman and a kart racer from @matthewmillerai. @joeldev_ said his five-hour usage limit “has barely budge[d].” These are first-day showcases, not controlled tests. community Watch the demos on X · Anthropic’s thread
07Where it falls short

- Hard, open-ended work: Anthropic says Opus 5.5 is still clearly stronger. vendor
- Facts: 54% accuracy on AA-Omniscience vs 66% for Opus 5.5, and about 6 points lower on Humanity’s Last Exam and SciCode. independent AA measured a lower hallucination rate than Opus (47% vs 59%), yet Anthropic’s system card says it “hallucinates more” on its own honesty evals. vendor
- More effort isn’t always better: on FrontierCode, Max scored below Xhigh (46.2% vs 52.1%). Anthropic blames timeouts and out-of-scope edits by code-review subagents. vendor
- Cyber fallback: some security tasks are handed to Sonnet 5; AA saw this on about 0.1% of its index tasks, mostly in Terminal-Bench. independent
08Price and how it stacks up

Pricing is unchanged from Sonnet 5: $2 per million input tokens, $10 per million output, $0.20 for cache reads and $2.50 for cache writes. It’s on the Claude Platform, AWS, Google Cloud and Microsoft Azure, and in GitHub Copilot and Devin. vendor
| Sonnet 5.5 | Opus 5.5 | GPT-6 Sol | GPT-6 Astra | |
|---|---|---|---|---|
| API price (in / out per 1M) | $2 / $10 | $4 / $20 | $2 / $10 | n/a |
| Output tokens per AA task (max) | ~193K | 119K | 31K | 27K |
| Terminal-Bench 4.0 (AA) | 64% | 60% | n/a | 60% |
| GDPval-AA (Elo) | 1844 | 1846 | 1487 | n/a |
Against GPT-6 Sol, at High effort the two are roughly tied on intelligence and cost per task. OpenAI’s models use far fewer tokens. independent
09Handle with care
- “Up to 30% less per task” (Anthropic) and “about 50% more” (AA) are both accurate; they refer to different effort settings.
- “50% cheaper than Opus” is true per token, not necessarily per task.
- AA ranked it #2 in its launch post; its model page later showed #3 of 216.
- Posts saying it’s free for everyone in Claude Code are unverified; the release notes don’t list plans.
- GPT-6 Sol’s GDPval-AA score may not yet reflect a fix for an image bug.
- VentureBeat reports that Google’s Gemini 3.8 Flash undercuts it on introductory pricing; that hasn’t been checked against Google’s own pricing.
- All figures are from launch day. AA tested a pre-release build with a structured-outputs bug and is re-running some evals.
10Sources
- vendor Anthropic: Introducing Claude Sonnet 5.5 · System card
- vendor @claudeai launch thread · early experiments · @bcherny
- independent Artificial Analysis thread · releases page · model page
- independent Zapier (@wadefoster) · Cognition · GitHub · VS Code
- press VentureBeat · The New Stack · SiliconANGLE
- community @haider1 · @neamtuz · @MatthewBerman · @matthewmillerai · @joeldev_