Claude Opus 5 tops the chart Anthropic built. It also loses the one it didn't.
Opus 5's launch numbers are a mix of three very different things: a benchmark Anthropic owns, a benchmark it trails on, and one genuinely impressive result an outside body actually checked. Only one of them is worth screenshotting — and it isn't the headline.

Image: Anthropic
Anthropic released Claude Opus 5 on the twenty-fourth of July, and the sentence that travelled fastest was the one about the lead: a new state of the art in coding and knowledge work, frontier-level intelligence at half the price, more than double the previous Claude on the flagship number. It is a good launch and, unusually for this beat, some of it survives a hard look. But 'some of it' is the operative phrase, and the way to see which parts survive is to do the boring thing the marketing never does — sort the numbers by who made the test and who checked the score. Do that, and the launch splits into three piles that deserve completely different amounts of trust.
Pile one: the chart Anthropic owns
The headline number is Frontier-Bench v0.1, where Opus 5 posts 43.3 percent, ahead of OpenAI's GPT-5.6 Sol at 34.4 and Anthropic's own larger model, Fable 5, at 33.7. It is a real gap and it is the number every write-up led with. It is also a benchmark Anthropic built, named, and is scoring itself first on, and the version string is 'v0.1' — which is to say it is brand new, and by definition has no track record of measuring anything yet.
This is not an accusation of cheating; it rarely is. It is the oldest structural problem in the field. When the company shipping the model also owns the test, the test measures, among other things, how much the company optimised for its own test. There is no held-out set you or I can inspect, no independent administration, no published methodology detailed enough to reproduce. The right response to a vendor topping its own freshly minted benchmark is not to call it false. It is to suspend judgment — to file the number under 'unverifiable as published' and wait for someone without a stake to run it. Compared to what? Compared to a test its maker also wrote. That is not zero information. It is just not the information the headline treats it as.
Two more launch numbers belong in this same pile, for the same reason. GDPval-AA, an Anthropic-run evaluation, gives Opus 5 an Elo of 1,861 against GPT-5.6 Sol's 1,736 — but it is human-graded, and the launch does not publish the rubric, the grader pool, or how self-preference was controlled when the lab grading the outputs is the lab that made one of the models. And Zapier's AutomationBench reportedly puts Opus 5 at 100 percent. A model scoring 100 on a benchmark has not proven it is perfect; it has proven the benchmark is saturated. A test everyone aces has stopped discriminating between the things it was built to tell apart. A ceiling is not a result.
Pile two: the chart it quietly lost
Here is the number that was in the materials and not in the headlines. On SWE-bench Pro — an independent software-engineering benchmark, not one Anthropic controls — Opus 5 scores 79.2 percent. Fable 5 scores 80.0. Mythos 5 scores 80.3. On the outside coding test, in other words, the new 'state of the art in coding' comes third, behind two other Anthropic models.
On the coding benchmark Anthropic didn't build, its new 'state of the art in coding' finishes third — behind two of its own models. That number was in the materials. It was not in the headline. — On SWE-bench Pro
Now, in fairness, 79.2 against 80.3 is almost certainly a statistical tie — the kind of gap that vanishes inside the margin of error on a benchmark of this size. But that cuts both ways, and it is precisely the point. If a one-point gap on SWE-bench Pro is noise, then a comparable gap anywhere else is noise too, including on the charts where Opus 5 is ahead. You cannot treat the differences that flatter the model as signal and the ones that don't as rounding. The honest reading of pile two is not 'Opus 5 is worse at coding.' It is that on the coding test the vendor didn't design, Opus 5 is roughly tied with the field — which is a very different claim from leading it, and it is the claim the launch chose not to make.
Pile three: the one result that actually counts
And then there is ARC-AGI-3, which is the reason this launch is genuinely interesting rather than merely loud. Opus 5 scored 30.2 percent, against a previous record of 7.8 held by GPT-5.6 Sol's maximum-effort configuration — close to four times the prior best. Anthropic frames it as roughly three times the next-best model. Either way it is an enormous margin, and here is why, uniquely, it survives the audit: ARC-AGI is run by the ARC Prize Foundation, an independent body, the test is explicitly designed to resist memorisation — you cannot have seen the answers in training — and, most importantly, on the twenty-fourth of July the ARC Prize Foundation published its own independently administered results confirming the score. That is the full set of conditions I spend most of my time complaining are missing: an outside test-maker, a memorisation-resistant design, and third-party verification of the number. When all three are present, credit is owed, and it is owed by name. This one counts.
The ARC Prize team went further and said the lead comes from genuinely better reasoning rather than test-taking tricks — which, from the people who built the test specifically to be hard to game, is about as strong an endorsement as this field produces. I want to be as careful crediting a result as I am puncturing one, so this is not a hedge: on the single most trustworthy benchmark in the launch, Opus 5 posted the most convincing frontier-reasoning result I have seen this year.
The caveats that remain are the ordinary ones, not disqualifying ones. Thirty percent is a low absolute score — the test is brutal, and Opus 5 still fails roughly seven of every ten tasks on it, which is worth remembering before anyone declares reasoning solved. ARC-AGI-3 is itself young, and young benchmarks move as more models pile in. But 'impressive and independently verified, on a hard test built to resist gaming' is a category that barely exists in AI launches, and this is a member of it.
The number nobody printed
Across the entire launch — every chart, every comparison, vendor-run and independent alike — there is one figure I could not find anywhere: a margin of error. No confidence intervals, no variance, no sample sizes disclosed for the benchmarks that matter. This is the field's chronic disease and Opus 5's launch has it as badly as any. It matters because without variance you cannot tell 43.3-versus-34.4 (probably a real gap) from 79.2-versus-80.3 (almost certainly not one). The two look identical on a slide — two bars, one taller — and they mean opposite things. A bar chart with no error bars is not a measurement. It is a picture of a measurement, and the picture is drawn by the party being measured.
You can see the whole problem in how differently the same statistical fact was treated depending on which way it pointed. The one-point SWE-bench Pro deficit was left out of the headlines; comparable one-point advantages elsewhere became 'state of the art.' Nobody published the interval that would tell you whether either gap was real. The reader is left to assume that the gaps flattering the model are significant and the ones that don't are noise — which is exactly backwards from how you would actually find out, and exactly the assumption a launch is built to produce.
What the numbers actually support
So strip the three piles down to what is left standing. The 'beats everyone' framing rests mostly on pile one — a benchmark Anthropic owns — and pile one is not evidence you can act on until an outsider runs it. Pile two shows the model roughly tied, not ahead, on the independent coding test, which quietly contradicts the 'state of the art in coding' line. And pile three contains one genuinely excellent, independently verified reasoning result on ARC-AGI-3 that deserves every bit of the attention the weaker claims are getting instead.
There is also a claim I have been ignoring because it is the one most likely to be true and least likely to be hyped: cost. Opus 5 lists at the same price as its predecessor while, by Anthropic's account, roughly matching its much larger sibling Fable 5 on several tasks at a fraction of the cost per task. That is an efficiency claim, not a supremacy claim, and efficiency claims are the ones this industry actually delivers on most reliably — they show up in your bill, which is the least gameable benchmark there is. If I were writing the launch, the honest headline would be two sentences long: Opus 5 is a large efficiency win and posts the year's most convincing independently verified reasoning result. It is not, on the evidence a person outside Anthropic can currently check, the across-the-board state of the art. Both of those are strong things to be. Only one of them is the thing the chart said, and it is not the one that owned the chart.
References
- Anthropic: Introducing Claude Opus 5
- SiliconANGLE: Anthropic launches Claude Opus 5 with efficiency, safety improvements
- The Decoder: Opus 5 blows past Fable 5 and GPT-5.6 Sol on the benchmark designed to measure real intelligence
- Digital Applied: ARC Prize verifies Claude Opus 5's ARC-AGI-3 record
- Vellum: Claude Opus 5 benchmarks explained
- OfficeChai: Claude Opus 5 scores 30% on ARC-AGI-3, triples previous best


