← Pooja Verma
BenchmarkedIssue · 28 Sep 2026
Model launch · about a 6-minute read

Claude Sonnet 5.5: near-Opus scores at half the token price

Independent tests confirm it sits two points behind Opus 5.5. The catch: at max effort it writes more output tokens per task than any model Artificial Analysis has measured.

Claude Sonnet 5.5 title over a view of Earth from Anthropic’s launch page

The short version

  • Near-Opus, confirmed: Artificial Analysis scores it 56 on its Intelligence Index, two points behind Opus 5.5’s 58 and #2 at launch. independent
  • Same price as Sonnet 5: $2 / $10 per million input/output tokens, half of Opus 5.5’s $4 / $20 per token. vendor
  • Token-hungry at max: about 193K output tokens per task, the most AA has measured, so a max-effort task costs about 50% more than with Sonnet 5. independent
  • Cheaper at lower effort: at Medium it beats Sonnet 5’s best Terminal-Bench score for less than a tenth of the cost per task, per Anthropic. vendor
  • Opus still leads on hard work: Anthropic itself says Opus 5.5 “remains clearly stronger” at complex, open-ended tasks. vendor

01What Anthropic announced

Sonnet 5.5 launched on 28 September as the second model in the Claude 5.5 family, six days after Opus 5.5. Anthropic pitches it as a faster, cheaper complement to Opus, “strongest at well-scoped everyday tasks, fixing bugs, and creating polished documents, slides, and spreadsheets.” vendor

02Anthropic’s scorecard

On Anthropic’s table, Sonnet 5.5 sits within a few points of Opus 5.5 almost everywhere and beats it on one row: Terminal-Bench 4.0 (multi-step command-line work), at 70.6% vs 66.4% for Opus at Xhigh effort. Sonnet 5 scored 10.3% on the same test. vendor

Anthropic’s benchmark table with the Sonnet 5.5 and Opus 5.5 columns highlighted
Anthropic’s own table, so read it as the vendor’s best case. GPT-6 Sol is missing from several rows. Source: Anthropic.

Elsewhere Opus keeps a small lead: CursorBench 4.0 (55.5% vs 57.8%), OSWorld 2.1 (80.1% vs 81.8%) and FrontierCode 1.1 (52.1% at Xhigh vs 54.4%). vendor

Sonnet 5’s Terminal-Bench score looks suspiciously low, but Artificial Analysis’s own run also shows a jump of about 50 points. independent

03What independent testers found

Artificial Analysis (AA) backs up the near-Opus claim: 56 on its Intelligence Index at max effort, two points behind Opus 5.5 and 18 ahead of Sonnet 5. independent

Artificial Analysis Intelligence Index bar chart with Claude Opus 5.5 at 58 and Claude Sonnet 5.5 at 56
Artificial Analysis Intelligence Index v4.3.2 at launch. The arrow marks the jump from Sonnet 5 (38) to Sonnet 5.5 (56). Source: Artificial Analysis.
TestResultKind
AA Intelligence Index (max)56, #2 at launch (Opus 5.5 58)independent
Terminal-Bench 4.0, AA’s run64% (Opus 5.5, GPT-6 Astra 60%)independent
GDPval-AA v2.1 (Elo)1844 (Opus 5.5 1846)independent
AA-Briefcase v1.1 (Elo)1811 (Opus 5.5 1822)independent
AA-Omniscience accuracy54% (Opus 5.5 66%)independent

04The catch: tokens

Price per token is not price per task. At max effort, AA measured about 193K output tokens per Intelligence Index task, the highest it has seen: roughly 60% more than Opus 5.5 or Sonnet 5 at max, and about 7× GPT-6 Astra. independent

Artificial Analysis post: Sonnet 5.5 uses about 193K output tokens per task, the most it has measured
Artificial Analysis on output tokens per task, with the Sonnet 5.5 (max) bar on the far right. Source: @ArtificialAnlys.

A max-effort task cost $7.60, about 50% more than with Sonnet 5. independent A chart of AA’s data from @haider1 puts Opus 5.5 (max) at 119K tokens, GPT-6 Sol at 31K and Astra at 27K. community

A widely shared claim of “150% more tokens” than Opus is wrong: AA’s figure is about 60%, and only at max effort, as @neamtuz noted. community

05Cheaper or pricier? It depends on effort

Anthropic says Sonnet 5.5 “costs up to 30% less per task” than Sonnet 5 vendor; AA says about 50% more independent. Both hold. They describe different effort settings.

Anthropic’s Terminal-Bench 4.0 accuracy-versus-cost chart, with Sonnet 5.5 at Medium circled
Terminal-Bench 4.0, score vs cost per attempt. Medium clears Sonnet 5’s best score for less than a tenth of the cost. Source: Anthropic.
EffortAA Intelligence IndexCost per AA task
Max56$7.60
Xhigh52$2.74
High47$1.08
Medium41$0.59
Low36$0.41

AA calls High the most competitive setting: very narrowly behind GPT-6 Sol on intelligence at effectively the same cost per task. independent

06Real use cases and demos

In a Claude Code bug-fix race posted by Anthropic’s Boris Cherny (median of 7 runs), Sonnet 5.5 finished in 12.3 s for $0.05 and 13.3K tokens, against 17.5 s, $0.08 and 18.2K for Sonnet 5. vendor

Boris Cherny’s post: Sonnet 5 and Sonnet 5.5 fixing failing tests side by side in Claude Code
Both models get the same repo and failing tests. Source: @bcherny (Anthropic). Watch the race on X.
Zapier AutomationBench chart: Sonnet 5.5 at 44.7% and $1.08 per task, above Opus 5.5 at 40% and $1.28
Zapier’s own AutomationBench; Zapier is a launch partner. Source: @wadefoster.

On Zapier’s test it scored 44.7% at Max vs 40% for Opus 5.5, at about $1.08 a task vs $1.28. It was strongest in Operations (59%) and weakest in HR (31.7%). independent

Builders posted one-shot games and sims, including a Fall Guys clone and a Starship sim from early tester @MatthewBerman and a kart racer from @matthewmillerai. @joeldev_ said his five-hour usage limit “has barely budge[d].” These are first-day showcases, not controlled tests. community Watch the demos on X · Anthropic’s thread

07Where it falls short

Anthropic: Opus 5.5 remains clearly stronger at complex, open-ended work requiring sustained judgment
Anthropic’s own caveat, from the launch post. Source: Anthropic.

08Price and how it stacks up

Pricing table: Sonnet 5.5 $2 input, $10 output; Opus 5.5 $4 input, $20 output
Price per million tokens. Source: Anthropic.

Pricing is unchanged from Sonnet 5: $2 per million input tokens, $10 per million output, $0.20 for cache reads and $2.50 for cache writes. It’s on the Claude Platform, AWS, Google Cloud and Microsoft Azure, and in GitHub Copilot and Devin. vendor

Sonnet 5.5Opus 5.5GPT-6 SolGPT-6 Astra
API price (in / out per 1M)$2 / $10$4 / $20$2 / $10n/a
Output tokens per AA task (max)~193K119K31K27K
Terminal-Bench 4.0 (AA)64%60%n/a60%
GDPval-AA (Elo)184418461487n/a

Against GPT-6 Sol, at High effort the two are roughly tied on intelligence and cost per task. OpenAI’s models use far fewer tokens. independent

09Handle with care

  • “Up to 30% less per task” (Anthropic) and “about 50% more” (AA) are both accurate; they refer to different effort settings.
  • “50% cheaper than Opus” is true per token, not necessarily per task.
  • AA ranked it #2 in its launch post; its model page later showed #3 of 216.
  • Posts saying it’s free for everyone in Claude Code are unverified; the release notes don’t list plans.
  • GPT-6 Sol’s GDPval-AA score may not yet reflect a fix for an image bug.
  • VentureBeat reports that Google’s Gemini 3.8 Flash undercuts it on introductory pricing; that hasn’t been checked against Google’s own pricing.
  • All figures are from launch day. AA tested a pre-release build with a structured-outputs bug and is re-running some evals.

10Sources

Benchmarked · the AI news, fact-checked · Facts as of 28 Sep 2026 · labels show who reported each number: vendor, independent, press, community