AnthropicProprietary

Claude Opus 5

Compare this model

Anthropic's Opus-tier flagship, released July 24, 2026. It topped the Artificial Analysis Intelligence Index at 61 on launch day while keeping Opus 4.8's $5/$25 pricing, and pairs a 1M-token context window with 96.0% on SWE-bench Verified and 90.4% on ARC-AGI-2.

Parameters

Undisclosed

Context Window

1M

License

Proprietary

Release Date

2026-07-24

Benchmark Performance

AA Intelligence Index

LMArena Elo

HLE

64.7

ARC-AGI-2

90.4

SWE-bench Verified

96.0

GPQA Diamond

93.2

MMLU-Pro

LiveCodeBench

AIME 2025

MATH-500

API Pricing

Input Price (per 1M tokens)

$5

Output Price (per 1M tokens)

$25

Billing Mode: standard

Strengths

  • Same $5/$25 price as Opus 4.8 with a large capability jump
  • Native self-verification; Frontier-Bench v0.1 more than doubles Opus 4.8
  • Leads on knowledge work (GDPval-AA v2 1,861 Elo) and science

Weaknesses

  • Anthropic's system card says it is not more capable overall than Fable 5
  • Loses DeepSWE v1.1 to GPT-5.6 Sol; trails Mythos 5 on cyber and legal work
  • Verbose and slow; breaking change - extended thinking cannot be disabled at xhigh or max

Use Cases

  • Long-horizon agentic coding and large-scale refactoring
  • Root-cause analysis of hard-to-reproduce bugs
  • Professional knowledge work and deep research

Deep Analysis

Artificial Analysis Intelligence Index

61

#1 overall at max effort; Claude Fable 5: 60, GPT-5.6 Sol: 59

SWE-bench Verified

96.0%

Anthropic-reported; vals.ai independently measured 97.0%

SWE-bench Pro

79.2%

SWE-bench Multilingual 89.5%; DeepSWE v1.1 68.8% (loses to GPT-5.6 Sol)

ARC-AGI-2

90.4%

ARC-AGI-1 97.5%; ARC-AGI-3 30.2%, roughly 4x the next-best model

Input / Output Price

$5 / $25 per 1M

Identical to Opus 4.8; half the cost of Fable 5 ($10/$50)

Context Window

1M tokens

128K max output; 300K via Batch API beta; knowledge cutoff May 2026

Strengths

  • Tops the Artificial Analysis Intelligence Index at 61 while charging exactly what Opus 4.8 charged ($5/$25 per 1M) - a capability jump with no price increase, and roughly half of Claude Fable 5.
  • Native self-verification. Opus 5 checks its own work without being told to, and Frontier-Bench v0.1 (43.3%) more than doubles Opus 4.8's 18.7% on long-horizon agentic coding.
  • Strong on knowledge work and science: 1,861 Elo on GDPval-AA v2 (over 100 points clear of Fable 5 and GPT-5.6 Sol), +10.2 points over Opus 4.8 on organic chemistry and +7.7 on protein tasks.

Weaknesses

  • Anthropic's own system card states Opus 5 "is not more capable overall than Claude Fable 5" - the index win is a composite result, not a clean sweep across every task.
  • It loses DeepSWE v1.1 to GPT-5.6 Sol by roughly four points, and Anthropic notes it still trails Mythos 5 on cybersecurity and on legal work.
  • Breaking API change: extended thinking is on by default and cannot be disabled at xhigh or max effort (a disabled-thinking request returns HTTP 400). Prompt scaffolds tuned for Opus 4.7/4.8 - especially "verify your work" instructions - can now degrade output by pushing the model toward over-checking.

Competitor Comparison

ModelArenaSWEGPQAPrice
Claude Fable 560 (AA Intelligence)80.3% (SWE-bench Pro)92.6%$10/$50 per 1M
GPT-5.6 Sol59 (AA Intelligence)64.6% (SWE-bench Pro)94.6%$5/$30 per 1M
Claude Opus 4.8N/A18.7% (Frontier-Bench v0.1)N/A$5/$25 per 1M

Claude Opus 5 is Anthropic's flagship Opus-tier model, released on July 24, 2026 as the fourth Claude 5 model in under two months (after Mythos 5, Fable 5 and Sonnet 5). The pitch is not raw capability but price-to-capability: Anthropic says Opus 5 "comes close to the frontier intelligence of Claude Fable 5 at half the price," and the API rate confirms it - $5 per million input tokens and $25 per million output, exactly what Opus 4.8 cost, against Fable 5's $10/$50. The model ID is claude-opus-5, the context window is 1M tokens with 128K max output, and it shipped same-day across the Claude API, Claude.ai, Claude Code and Claude Cowork. It is the default on Claude Max and the strongest option on Claude Pro.

Independent testing backed the launch claims immediately. On July 24, Artificial Analysis placed Opus 5 at #1 on both its Intelligence Index (61, the highest score it had ever recorded) and its Agentic Index (55.3), with a joint-first 78 on the Coding Index. Anthropic's own published figures are equally aggressive: 96.0% SWE-bench Verified, 79.2% SWE-bench Pro, 89.5% SWE-bench Multilingual, 90.4% ARC-AGI-2, 97.5% ARC-AGI-1, 93.2% GPQA Diamond, 70.6% OSWorld 2.0, 90.8% BrowseComp and a perfect 42/42 on IMO 2026. The two results worth singling out are Frontier-Bench v0.1 (43.3% against Opus 4.8's 18.7%) and ARC-AGI-3 (30.2% against GPT-5.6 Sol's 7.8%), both of which measure the ability to keep working after the obvious approach fails.

The honest caveat is that Anthropic does not actually call Opus 5 its most capable model. The system card says plainly that Opus 5 "is not more capable overall than our most capable general-access model, Claude Fable 5," and Anthropic still recommends Fable 5 for work a model might run autonomously for days. Both statements can be true at once: Fable 5 remains the nominal ceiling, while Opus 5 tops the leading independent composite index and wins several agentic and coding evaluations at half the token cost. For most teams the routing decision is straightforward - Opus 5 becomes the everyday default, and Fable 5 gets reserved for the genuinely hardest jobs.

Analysis generated: 2026-08-11