← Pooja Verma
BenchmarkedAll issues

Benchmarked

The AI news, fact-checked. Every number is labelled by who reported it: the company, an independent tester, the press, or the community.

3 Oct 2026 · Weekly roundup · Sep 27 – Oct 3, 2026 · about a 7-minute read

The week Washington stepped into AI

A voluntary White House accord, a reported FTC probe, Anthropic’s leaked IPO numbers and a turbulent week at OpenAI, plus where the model race stands.

1 Oct 2026 · Rumor check · about a 6-minute read

Is Claude Fable 5.5 already here?

Anthropic hasn’t announced it. Some users say chats labelled Fable 5.1 are behaving like a newer model, but the name, the quality and the release date are all unconfirmed.

30 Sep 2026 · Product launch · about a 6-minute read

Meta’s Muse: the most polished consumer agent yet, after a rough first three weeks

A free tier took Muse to #1 on both app stores. The model is mid-pack, and launch brought an Amazon block, a privacy dispute and a patched Mac flaw.

30 Sep 2026 · Model launch · about a 6-minute read

Gemini 4 Argon: is Google back on top?

Google's own chart says it wins 12 of 18 benchmarks. Independent testers say it's top-tier, but not the best, and almost nobody can use it yet.

29 Sep 2026 · Event recap · about a 6-minute read

OpenAI DevDay 2026: three launches that matter, and a catch

Always-on agents, a cheap near-flagship model and an up-to-8× speed tier led more than 20 announcements. Independent tests call the model a value pick, not the smartest, and the $200 plan shrank.

28 Sep 2026 · Model launch · about a 6-minute read

Claude Sonnet 5.5: near-Opus scores at half the token price

Independent tests confirm it sits two points behind Opus 5.5. The catch: at max effort it writes more output tokens per task than any model Artificial Analysis has measured.

25 Sep 2026 · Product launch · about a 6-minute read

Jev: the AI model that can’t write a sentence

TypeSafe AI’s new model only makes typed decisions, fast and very cheaply. Early independent tests back the speed and cost, but its accuracy looks like a good small model’s, and it always answers, even when it shouldn’t.

24 Sep 2026 · Model launch · about a 6-minute read

Claude Opus 5.5: a narrow lead at a lower price

Independent leaderboards put Anthropic’s new mid-priced model first, but the margin shrinks outside Anthropic’s own tests, and at max effort the token bill climbs fast.

Benchmarked · 8 issues · 24 Sep – 3 Oct 2026