StartClaude Opus 5 Just Edged Out Fable 5 on the Benchmarks. Almost Nobody Noticed.
By Addy · July 25, 2026
Claude Opus 5 launched on July 24 with a result that should have reset the frontier conversation. In the July 24–25 Artificial Analysis snapshot, it scored 61 on the Intelligence Index, one point ahead of Claude Fable 5 at 60 and two ahead of GPT-5.6 Sol at 59. The lead is narrow. It is also real: Opus 5 became the highest-scoring model in the snapshot while costing less per measured task than Fable 5.
Anthropic priced Opus 5 at USD 5 per million input tokens and USD 25 per million output tokens, the same rate card as Claude Opus 4.8 and half of Fable 5’s token price. Artificial Analysis measured Opus 5 at USD 2.03 per Intelligence Index task against Fable 5 at USD 2.75, a 26% gap. That is a meaningful frontier result, especially for teams already paying for the larger model.
It still did not produce the kind of launch reaction Anthropic’s previous Opus releases received. The reason is not that the model failed. The more interesting explanation is that Opus 5 arrived after the surprise had already leaked, after Anthropic had shipped three other major models in less than eight weeks, and after early users had started arguing about whether the new model was actually pleasant to work with.
The benchmarks were never the bottleneck. The timing was.
The Benchmarks Actually Hold Up
The Intelligence Index lead is only one point, and the leaderboard is not a clean sweep. Fable 5 remains ahead of Opus 5 on SWE-bench Pro in the comparison published around launch, 80.0% to 79.2%. That is exactly the kind of result that gets lost when a composite score is turned into a winner-takes-all headline.
The broader coding picture still favors Opus 5. Artificial Analysis measured it at 96% on SWE-bench Verified against Fable 5 at 95%, and its Coding Agent Index placed Opus 5 at the top or joint top depending on the effort configuration. On GDPval-AA v2, its 1,861 Elo score sits more than 100 points above Fable 5. On AA-Briefcase, a professional knowledge-work evaluation run through an open agent harness, it leads Fable 5 by 146 Elo.
Anthropic’s own release evidence points in the same direction. On CursorBench 3.2 at maximum effort, the company says Opus 5 comes within 0.5% of Fable 5’s peak score at half the cost per task. On Frontier-Bench v0.1, Anthropic says Opus 5 more than doubles Opus 4.8’s performance at a lower task cost. Those are vendor-run results, and the footnote matters: Opus 4.8 served as a safety fallback for some refused tasks. The numbers are still important, but they are not interchangeable with an independent leaderboard.
The benchmark pattern is therefore unusually coherent. Opus 5 wins the broad composite, wins or ties the agentic coding layer, wins the knowledge-work measures, and stays close to Fable 5 on the coding test most likely to shape developer decisions. It loses some individual tests. That is what a real frontier result looks like when the models are close enough for the task definition to decide the winner.
The exception is cybersecurity. Anthropic says Opus 5 was not trained specifically on cyber tasks and remains behind Mythos 5 on exploit development. On ExploitBench, the system-card comparison puts Opus 5 at 99 full working exploits against Mythos 5 at 132. Opus 5 can identify vulnerabilities at a similar level, but its safeguards intervene much more heavily when the task turns from finding a flaw into turning it into a usable attack.
That distinction matters because Anthropic’s public model family is now deliberately split. Fable 5 is the generally available frontier model with strong restrictions. Opus 5 is the broadly useful model with fewer restrictions on legitimate research but tighter limits on dangerous cyber work. Mythos 5 remains the model Anthropic does not want most users to access at all. The capability ladder is real, but so is the policy ladder.
Effort Is the Product Feature Nobody Is Counting
Opus 5 does not behave like one fixed model. Anthropic exposes five effort levels: low, medium, high, xhigh, and max. The setting changes how many tokens the model spends, how long it works, and how much the finished task costs.
Artificial Analysis measured roughly an eightfold range in output-token usage across effort settings on GDPval-AA. That is not a cosmetic control. It means one model can act like a fast workhorse for a simple request and a slow, expensive project team for an ambiguous one. The model name stays the same while the economics change underneath it.
This is where Opus 5’s cost story becomes more useful than its headline price. At max effort it can approach Fable-level work with a lower task bill. At lower effort it can give up some peak performance and preserve budget for the next turn. The customer is not choosing only between Opus and Fable. They are choosing how much reasoning to buy for each job.
The tradeoff is also a warning. A benchmark run at max effort says almost nothing about what a production deployment will cost if users select low, medium, or high. The reverse is true too: a low-effort result should not be treated as evidence that the underlying model lacks frontier capability. Opus 5 is a family of cost-performance points wearing one name.
The Reveal Had Already Happened
The first reason the launch felt quiet is simple: the model had already been discussed for weeks. A codename, Honeycomb, appeared in a Cursor model-picker screenshot around July 8. The tooltip described a one-million-token context window, per-turn controls, safety fallbacks, and an extra-high effort mode. A separate model listing later appeared in Google Cloud’s catalog. Neither artifact was a supported public API contract, but both gave the rumor production vocabulary.
That changed the emotional shape of the announcement. Anthropic was not introducing an unknown model to a waiting audience. It was confirming a model that parts of the developer internet had already imagined, compared, and argued about. The official post supplied the benchmark table. It did not supply surprise.
This is a different kind of leak from a vague claim that a new model exists. The screenshots described controls developers could recognize. They made the rumored model feel operational before anybody could actually call it. Once a launch becomes a confirmation, even a strong result arrives with some of its attention already spent.
The Fourth Anthropic Release in Eight Weeks
The timing did the rest. Mythos 5 and Fable 5 arrived together on June 9. Sonnet 5 followed on June 30. Opus 5 landed on July 24. Anthropic shipped four major model releases in a span short enough that each announcement had to compete with the memory of the previous one.
That density changes how improvement feels. A model that would have looked like a generational event in a quiet quarter can look like the next item in a product calendar when users have just learned a new model’s pricing, fallback rules, safety behavior, and API identifier. Developers do not reset their expectations after every launch. They carry the last few disappointments forward.
The release sequence also made the product line harder to explain. Fable 5 is the public flagship. Mythos 5 is the restricted model. Sonnet 5 is the cheaper general option. Opus 5 is now the model that sits below Fable in name but above it on the current composite, while remaining cheaper at the token level. Anthropic’s lineup has become more capable and less intuitive at the same time.
The Workflow Reaction Was Mixed for a Reason
The most useful early reaction did not come from a leaderboard. Every spent its first week with Opus 5 across coding, writing, and agentic workflows, and described it as brilliant in flashes and frustrating in practice. It argued with instructions, stopped before finishing, and did not cooperate cleanly with the team’s existing skills and plugins.
That is a more important finding than a bad first impression might suggest. A frontier model is not used in a blank chat window. It is dropped into prompts, tools, memory systems, permission layers, test harnesses, and habits built around its predecessor. If those assumptions no longer fit, a model can be objectively stronger and subjectively worse on the first day.
Every’s account also reported that results improved after the team rebuilt its setup instead of carrying old workflows forward unchanged. That is a mixed result, not a dismissal. Opus 5 may require a new operating style precisely because it is more willing to verify, push back, and spend effort before it commits. The same judgment that helps on a long debugging task can feel like friction when a user expects instant compliance.
This is the gap between a benchmark win and a product win. Benchmarks ask whether a model can solve a fixed task. Workflows ask whether a team can predict when it will solve the task, how much supervision it needs, and whether its extra care is worth the time it takes.
The Benchmarks Were Never the Bottleneck
Claude Opus 5 is a strong result by the numbers that matter. It narrowly leads the July Artificial Analysis snapshot, matches or beats Fable 5 on most of the coding and knowledge-work evidence, and does so at half Fable’s token price. It also gives customers an effort control that makes the cost-performance curve unusually explicit.
Its weaknesses are equally instructive. It is not ahead on every coding benchmark. It remains behind Mythos 5 on dangerous cyber work by design. Its best independent composite score is only one point above Fable 5, and live leaderboard pages can change as evaluations and index versions update. A one-point lead is evidence of parity at the frontier, not proof of permanent dominance.
The muted response says more about the release environment than about the model. Honeycomb made the result feel late before Opus 5 shipped. Four Anthropic launches in eight weeks made another model feel routine. Mixed early workflow reports gave developers a reason to wait before rebuilding their systems. None of those facts cancels the benchmark result. They explain why the result did not travel.
Anthropic built the model. Artificial Analysis measured the lead. The internet had already moved on to the next rumor.
That is the honest reading of Opus 5: a quiet number-one finish, a cheaper path to near-Fable capability, and a launch whose timing made an unusually strong model look like another Thursday update.
Sources:
- Introducing Claude Opus 5 — Anthropic
- Opus 5: Fable 5 level intelligence at a lower cost per task — Artificial Analysis
- Vibe Check: Claude Opus 5 Is Brilliant in Flashes, Frustrating in Practice — Every
Previously on TheQuery: