The latest foundation model developed by Anthropic. It achieves significantly improved performance and efficiency compared to the previous generation.
Parameters
Undisclosed
Context Window
1M
License
Proprietary
Release Date
2026-05-28
Japanese Language Capability
Multilingual model with strong Japanese language processing capabilities.
API Pricing
Input Price (per 1M tokens)
$5
Output Price (per 1M tokens)
$
Billing Mode: standard
Strengths
Weaknesses
Use Cases
Deep Analysis
Arena Elo
1477
BenchLM overall; 1537 on coding subset
SWE-bench Pro
69.2%
Field-leading; GPT-5.5: 58.6%, Gemini 3.1 Pro: 54.2%
SWE-bench Verified
88.6%
Up from 87.6% on Opus 4.7
GDPval-AA
1,890 Elo
67% implied win rate vs GPT-5.5
Intelligence Index
61.4
Was #1 at release; +4.1 over Opus 4.7
Input / Output Price
$5 / $25 per 1M
Unchanged from Opus 4.7; fast mode $10/$50
Strengths
- ・Field-leading agentic coding performance (SWE-bench Pro 69.2%) with significant margin over GPT-5.5 and Gemini 3.1 Pro
- ・~4x reduction in letting code flaws pass without flagging, improving reliability in long-horizon agentic workflows
- ・Strongest long-context reasoning at 1M tokens (GraphWalks BFS 68.1% vs GPT-5.5's 45.4%) and leading knowledge-work Elo (GDPval-AA 1,890)
Weaknesses
- ・GPT-5.5 still leads on terminal-based DevOps coding (Terminal-Bench 78.2% vs 74.6%) and uses ~30% fewer turns for agentic tasks
- ・Prompt injection resistance regressed from ~2.3% to ~7% attack success rate, a concern for adversarial environments
- ・High token consumption and output pricing ($25/M) makes it expensive vs Gemini 3.1 Pro ($12/M output) for cost-sensitive high-volume workloads
Competitor Comparison
| Model | Arena | GPQA | Price |
|---|---|---|---|
| GPT-5.5 (xhigh) | ~1480* | ~92%* | $25/$30 |
| Gemini 3.1 Pro | ~1460* | ~92%* | $2/$12 |
| Claude Fable 5 | N/A | N/A | $10/$50 |
Claude Opus 4.8, released May 28, 2026, is Anthropic's incremental but substantive upgrade to the Opus 4 line. It builds on Opus 4.7 with meaningful gains in agentic coding (SWE-bench Pro from 64.3% to 69.2%), knowledge work (GDPval-AA from 1,753 to 1,890 Elo), and scientific reasoning (Humanity's Last Exam up ~3 points), while maintaining the same $5/$25 pricing and 1M-token context window. At launch it claimed the #1 spot on the Artificial Analysis Intelligence Index at 61.4, leapfrogging GPT-5.5 by 1.2 points.
Beyond benchmark numbers, the release's most consequential improvements are behavioral: Opus 4.8 is approximately four times less likely than its predecessor to let code flaws pass unremarked, and early testers consistently report better judgment, more appropriate pushback on flawed plans, and stronger self-correction in agentic workflows. The model also introduces Dynamic Workflows (a research preview enabling up to 1,000 parallel subagents in Claude Code), Effort Control (a 5-level dial for reasoning depth), and a new Fast Mode at $10/$50 offering 2.5x throughput. Anthropic positioned Opus 4.8 as the dependable production tier beneath the newly announced Fable 5 and Mythos-class models, which handle even more demanding tasks at higher price points.
The model's primary weaknesses are in areas where competitors retain structural advantages: GPT-5.5 leads on terminal-based coding tasks and achieves higher turn efficiency on agentic benchmarks, Gemini 3.1 Pro is roughly 2x cheaper on output tokens for cost-sensitive workloads, and prompt injection resistance regressed from the prior version. Creative writing remains a lateral move rather than an improvement. Nevertheless, for teams building agentic coding pipelines, code review automation, or long-context analysis tools, Opus 4.8 represents the strongest mainstream-tier option as of mid-2026.
Sources
- Introducing Claude Opus 4.8 — Anthropic
- Claude Opus 4.8 Takes the Lead on the Artificial Analysis Intelligence Index
- Claude Opus 4.8 Benchmark Scores & Evals — BenchmarkList
- Claude Opus 4.8 Benchmarks, Pricing & Speed — BenchLM.ai
- Claude Opus 4.8 Review: Better At What It's Good At — Tech Yahoo
- 5 Biggest Upgrades: Claude Opus 4.8 vs 4.7 — Android Authority
- Claude Opus 4.8 Review: Reliability Over Raw Scores — Awesome Agents
- Claude Opus 4.8 Review — AI Panel Verdict — TopReviewed.ai
- Claude Opus 4.8 vs Opus 4.7: What Actually Changed — Contra Collective
- Claude Opus 4.8 Benchmarks, Scores, Rankings — AllThings.how
Analysis generated: 2026-07-17