Frontline Hotspot
Frontline Hotspot

Claude Opus 5: Topping Both the Intelligence and Agent Leaderboards

Anthropic released Claude Opus 5 on July 24, 2026, topping both the Artificial Analysis Intelligence Index (61) and Agentic Index (55.3), priced at $5/$25 unchanged from Opus 4.8. Its real edge is agentic capability (leading AA-Briefcase/GDPval), but it runs slow, burns tokens, and its hallucination rate rose to 50%; it is built for complex long tasks, not everyday Q&A.

Published July 31, 20265 min read
<!-- claude-opus-5-hotspot | hotspot | Claude Opus 5: Topping Both the Intelligence and Agent Leaderboards -->

On July 24, 2026, Anthropic released Claude Opus 5 - its fourth model in under two months, after Mythos 5, Fable 5, and Sonnet 5. On Artificial Analysis's boards it took two firsts at once: 61 on the Intelligence Index and 55.3 on the Agentic Index. Most models split these two lines - the smart ones aren't always the ones that get work done, and vice versa. Opus 5 stacking both is the real signal of this release. But "topping" needs caveats, spelled out below.

What "Top of Both Boards" Actually Means

The numbers first. On the Artificial Analysis Intelligence Index, Opus 5 (max effort) scores 61, in first place, ahead of Claude Fable 5 at 60, GPT-5.6 Sol at 59, Kimi K3 at 57, and the previous Opus 4.8 at 56. On the Agentic Index, Opus 5 leads at 55.3, with GPT-5.6 Sol second at 54.0 and Fable 5 third at 52.8. Further down, GLM-5.2 sits at just 43.1 and Gemini 3.6 Flash at 38.7, clearly behind the front pack.

To read these scores you first need to know how they are built. The Intelligence Index normalizes several benchmarks into one composite; the Agentic Index is assembled from agentic evaluations like GDPval-AA v2 and 𝜏³-Banking, measuring whether a model, once wired into an agent harness, can independently finish a professional task. Opus 5 also has five effort settings (low, medium, high, xhigh, max): the higher the setting, the more tokens and turns, and the higher the score. Every figure above is marked max effort - the model's top possible output.

But "first" here comes with a discount. First, 61 is only one point above Fable 5; Artificial Analysis itself calls Opus 5 "narrowly the most intelligent" and "effectively tied" with Fable 5 - not a clean break. Second, both scores are benchmark-run, not human-voted. As of 2026-07-31, Opus 5 has not yet entered the Arena or Scale SEAL vote-based boards, so "topping" refers to the benchmark boards, not the reputation boards. These boards move in real time, so don't treat today's 61 as the final word.

Agentic Capability Is the Real Headline

More than the Intelligence Index, the agentic capability is where Opus 5 actually separates itself. On AA-Briefcase (agentic knowledge work), Opus 5 scores 1720 Elo - 146 above Fable 5 - and on GDPval-AA v2 it scores 1861 Elo, 114 above Fable 5. It hits 89% on Terminal-Bench v2.1 (terminal tasks), and on the Coding Agent Index, Opus 5 (xhigh) paired with Claude Code sits in a tie for first. It is also strong on scientific reasoning: 30% on ARC-AGI 3 versus 8% for the next-best model - nearly a 4x gap, independently confirmed by ARC Prize.

Anthropic's own description is "thoughtful and proactive." In his July 24 blog post, Simon Willison noted that Opus 5 is "relentlessly proactive" - it keeps driving a task forward instead of waiting for instructions - and singled out one detail: Opus 5 is Anthropic's "least prompt injectable model yet." For teams wiring the model into external tool chains that read untrusted web pages, that line is worth more than any benchmark score.

But agentic capability has a cost. On AA-Briefcase tasks, Opus 5 at max effort averages 36.2 minutes and 103 turns, about 50% longer than Opus 4.8's 24.1 minutes and 55 turns. In other words, it "can do long work," but slowly, burning tokens. One concrete contrast: from low to max effort, GDPval-AA v2 spans 407 Elo and roughly 8x in token usage - the same model on max and on low is two very different cost curves.

There is also an honest regression: its hallucination rate rose 14 points over Opus 4.8 to 50% - the model is more willing to answer even when uncertain - and its factual knowledge (AA-Omniscience) still trails Fable 5. It is not strong on every subject either: on CritPt, a frontier physics eval, it sits behind GPT-5.6 Sol and Terra. So don't treat it as an encyclopedia; its strength is "finishing a whole task," not "getting every fact right."

Where the Rivals Stand

Place Opus 5 back in the competitive field and its position is "first on both boards, but with a close challenger on each." On the Intelligence Index, GPT-5.6 Sol sits just 2 points back at 59, and at $2/$8 per million tokens it is far cheaper per token than Opus 5's $5/$25. Fable 5 is Anthropic's own stronger flagship, 1 point back at 60, but priced at $10/$50 - double Opus 5. Opus 5's pitch is exactly this: at half of Fable 5's token price, it approaches Fable 5's intelligence and actually beats it on agentic work.

The real contest is Opus 5, Fable 5, and GPT-5.6 Sol wrestling in the front row, with Gemini and GLM still in the back. On the coding line, be precise: GPT-5.6 Sol with Codex leads the Coding Agent Index alone at 80, while Opus 5 (xhigh) with Claude Code is part of a "tie for first"; Terminal-Bench's 89% merely matches GPT-5.6 Sol, not beats it. So on coding Opus 5 has nothing like its intelligence and agent lead - it is a close-quarters fight with GPT-5.6.

The decisive variable is cost per completed task, not per-token price. Artificial Analysis gives comparable numbers: average cost per Intelligence Index task is $2.03 for Opus 5 (max), $2.75 for Fable 5 (with fallback), $1.80 for Opus 4.8 (max), and $1.53 for Sonnet 5 (max). On raw token price Opus 5 is expensive, but at high/xhigh effort it beats Opus 4.8 and Sonnet 5 on cost per task; at lower effort, the GPT-5.6 family sits ahead on the intelligence-versus-cost frontier. Opus 5 on simple tasks is waste; its home ground is complex reasoning and long agentic work. One cost lever worth pulling: cache hits cost $0.50 (90% off) and the Batch API halves prices - always enable these for long contexts and bulk eval runs.

What It Means for Developers and Everyone Else

For developers, the upshot is that the reliable ceiling on long agent tasks has lifted again. If you are running Claude Code, multi-step code maintenance, or long knowledge-work flows, it is worth switching up to try it. It has a 1-million-token context window, up to 128K output tokens, and a May 2026 knowledge cutoff. Pricing is $5/$25 (input/output per million tokens), identical to Opus 4.8; a Fast mode at $10/$50 runs about 2.5x quicker for when you are in a hurry. But measure cost-per-task on your own workloads - the per-token price misleads.

There is also an often-overlooked dividend: safety. Anthropic lists Opus 5 as its "most capable generally available model for scientific research," and biology-related requests blocked on Fable 5 now route to Opus 5. Combined with it being the least prompt-injectable generation yet, these two points directly decide whether long-running, unattended agents that ingest external data can ship at all.

For everyone else, Opus 5 is already in Claude subscriptions (Pro from $20/month, Max $100/$200). For everyday Q&A and writing, Sonnet 5 ($2/$10) or even Haiku 4.5 ($1/$5) is enough; Opus 5 is built for the kind of job where you need it to run for half an hour and finish a whole task on its own. The Claude 5 generation now spans four tiers - Haiku, Sonnet, Opus, Fable - from cheapest to strongest flagship. The test is simple: does this job need it to loop through dozens of turns by itself? If yes, use Opus 5; if not, step down a tier and save the money.

A new model "topping both boards" sounds imposing, but it comes down to one thing on the ground: can it, on your hardest, most time-consuming task, make fewer mistakes and take fewer back-and-forth rounds than the last generation. If yes, the upgrade is worth it. What you should actually do is not refresh the leaderboard but take your most painful task, run it a few times each on Opus 5 and Opus 4.8, and compare error rate and round count - that is the only way to translate a benchmark into your own workload.


References

This article is AI-assisted and human-edited. Last updated: 2026-07-31

Related