Benchmarked
The AI news, fact-checked. Every number is labelled by who reported it: the company, an independent tester, the press, or the community.
The week Washington stepped into AI
A voluntary White House accord, a reported FTC probe, Anthropic’s leaked IPO numbers and a turbulent week at OpenAI, plus where the model race stands.
Is Claude Fable 5.5 already here?
Anthropic hasn’t announced it. Some users say chats labelled Fable 5.1 are behaving like a newer model, but the name, the quality and the release date are all unconfirmed.
Meta’s Muse: the most polished consumer agent yet, after a rough first three weeks
A free tier took Muse to #1 on both app stores. The model is mid-pack, and launch brought an Amazon block, a privacy dispute and a patched Mac flaw.
Gemini 4 Argon: is Google back on top?
Google's own chart says it wins 12 of 18 benchmarks. Independent testers say it's top-tier, but not the best, and almost nobody can use it yet.
OpenAI DevDay 2026: three launches that matter, and a catch
Always-on agents, a cheap near-flagship model and an up-to-8× speed tier led more than 20 announcements. Independent tests call the model a value pick, not the smartest, and the $200 plan shrank.
Claude Sonnet 5.5: near-Opus scores at half the token price
Independent tests confirm it sits two points behind Opus 5.5. The catch: at max effort it writes more output tokens per task than any model Artificial Analysis has measured.
Jev: the AI model that can’t write a sentence
TypeSafe AI’s new model only makes typed decisions, fast and very cheaply. Early independent tests back the speed and cost, but its accuracy looks like a good small model’s, and it always answers, even when it shouldn’t.
Claude Opus 5.5: a narrow lead at a lower price
Independent leaderboards put Anthropic’s new mid-priced model first, but the margin shrinks outside Anthropic’s own tests, and at max effort the token bill climbs fast.