Claude Opus 5 Benchmarks: What the 63.05 Actually Means

The Claude Opus 5 API everyone is quoting reduce to one headline figure right now: 63.05 on the Artificial Analysis Intelligence Index, the current top score out of 185 models — a single independent reading of one composite index, at one effort setting, on one day, not a fixed property of the model. This piece is the plain-language version of what that number means and how much weight to give it; Claude Opus 5 carries the same 63.05 and get re-checked whenever the board moves.

The fastest way to misread a benchmark is to treat it like a spec sheet. A model does not have an IQ; it has a score on a particular test, run at a particular reasoning effort, on a particular harness, on a particular day. Claude Opus 5 makes that point louder than most, because the same model currently holds four different scores on the same index depending on how much thinking you let it do. Anthropic released Opus 5 on July 24, 2026, as its flagship reasoning model, and Artificial Analysis positions it as “near-frontier at half the price of Claude Fable 5.” Here is how to read the scores behind that positioning.

What the 63.05 actually is

The 63.05 is the max-effort configuration of the Intelligence Index. Artificial Analysis runs a fixed composite of reasoning, knowledge, maths and coding tasks across every model it tracks — the closest thing to an apples-to-apples general index available — and it labels this entry “Claude Opus 5 (Adaptive Reasoning, Max Effort).” On its live board, which we checked August 22, 2026, that configuration scores 63.05, first of 185 models.

The ranking context, all from the same independent board: Claude Fable 5 (with fallback) sits at 62.07, GPT-5.6 Sol (max) at 60.93, Grok 4.6 (high) at 60.92, and GPT-5.6 Luna (max) at 52.32. Opus 5 leads its nearest rival by a little over a point — a meaningful but narrow margin that could flip on any re-run. That is the headline to quote, and also the reason not to over-weight it.

The effort ladder: one model, four scores

The headline only makes sense with the configuration attached. Anthropic’s “Adaptive Reasoning” lets you trade thinking time for quality, and Artificial Analysis publishes the same model at four rungs on the same index: max 63.05, xhigh 62.52, high 61.48, and medium 58.64. That is a 4.4-point swing entirely within one model, driven purely by how much reasoning budget you request.

This matters in every comparison you will read. If one article quotes Opus 5 at max effort and a rival at default effort, you are not comparing like for like. The widely quoted number — 63.05 — is the most expensive configuration to run, and it is the one to use only when you actually need the quality ceiling.

Claude Opus 5

Where it wins and where it loses

Start with the wins. On the same Artificial Analysis board, Claude Opus 5 (max) posts the highest Intelligence Index at 63.05 and the highest Omniscience Index at 37.07. But the omniscience caveat matters: Claude Fable 5 leads that specific index at 43.3, so “highest omniscience” means highest among the rest, not highest overall.

Then the trade-offs. Output is slow: Opus 5’s median output speed, per Artificial Analysis, is 61.8 tok/s, versus 156.6 for GPT-5.6 Luna and 399 for Gemini 3.7 Flash on the same board. Running the full index costs $2.34 per task, and at its list prices of $5 in / $25 out per 1M tokens (with cached input at $0.50, per Anthropic’s current API list) it ranks only #74 of 183 on cost. Production telemetry agrees: OrcaRouter’s own 7-day data shows a p50 time-to-first-token of 7.34 s (p95 10.00 s) for Opus 5, against 1.33 s for GPT-5.6 Luna. Opus 5 thinks longer before answering — it is an output-quality model, not a latency model.

Metric Claude Opus 5 (max) Claude Fable 5 (with fallback) GPT-5.6 Sol (max)
Intelligence Index 63.05 — #1 of 185 62.07 60.93
Omniscience Index 37.07 43.3 —
Median output speed 61.8 tok/s — 73.7 tok/s
Cost per index task $2.34 — $1.23
List price (in/out per 1M) $5 / $25 $10 / $50 $5 / $30

 

Sources: Artificial Analysis live board for intelligence, omniscience, speed and per-task cost (checked 2026-08-22); Anthropic and OpenAI API list prices (GPT-5.6 shown after OpenAI’s price cut); OrcaRouter telemetry for production latency.

Claude Opus 5

The honest caveats: one index, one scaffold, silent re-scores

Three caveats change how you should read 63.05:

  • One index. 63.05 is a single composite index — Artificial Analysis’ Intelligence Index, an independent measurement. Anthropic does not publish an official benchmark for Opus 5 that this derives from, so treat it as third-party data, not a vendor claim.
  • One scaffold. The index runs every model on the same fixed harness, which is exactly what makes it comparable — but it also means the number describes that harness, not every way to run the model. A different scaffold with the same checkpoint produces a different number, and that difference is real, not noise.
  • Silent re-scores. Quote the live board: the current page at artificialanalysis.ai/models/claude-opus-5, checked August 22, 2026, reads 63.05 — and it has read differently in the past. The board re-scores silently as the task mix and baselines shift; earlier snapshots had Opus 5 at 63. If you see a near-miss figure in an older article, you are looking at an older snapshot, not a different model. Always date the number you quote.

Why the same model scores differently across harnesses

The same model landing at different scores in different places is not a glitch — it is how benchmarks work. Every published figure is a product of four variables: the task set, the harness it runs in, the scaffold or sampling around the model, and the reasoning effort. Change any one and the number moves. That is why Opus 5 can top AA’s composite index yet drop 4.4 points on the same index when run at medium effort, and why a vendor-run comparison in which every model uses a different scaffold is a product comparison, not a model comparison. When someone quotes a benchmark at you, ask four questions: who ran it (vendor or independent), on what harness, at what effort, and on what date.

The practical move is to stop treating public benchmarks as a purchasing decision and start treating them as a shortlist filter. Once a model is on your shortlist, the only benchmark that settles anything is your own task set — and that is easier to run when the candidates sit behind one API. OrcaRouter carries Claude Opus 5 among 200-plus models at 0% markup, so testing effort levels or A/B-ing against Claude Fable 5 is a config change rather than a procurement cycle.

The takeaway

Claude Opus 5 is the current #1 on the Artificial Analysis Intelligence Index at max effort — a real, independently measured result, and the number to quote is 63.05 with the configuration and the date attached. But the lead is narrow (Claude Fable 5 sits at 62.07), Fable 5 leads omniscience (43.3 vs 37.07), the model is slow (61.8 tok/s output), and it holds the top slot only at max effort. Buy it for quality-critical, latency-tolerant work; drop its effort level or pick a faster model when speed or budget wins; and never compare it against anything else without matching the effort setting on both sides.

Leave a Comment