AI Model Rankings
Comprehensive AI model rankings across 20 benchmarks. Detailed comparisons by category.
Comprehensive Ranking
Overall AI model ranking across HLE, ARC-AGI-2, FrontierMath, SWE-bench Verified, and τ²-Bench.
5 benchmarks
Coding Capability
Programming ability benchmarks: SWE-bench Verified, LiveCodeBench, SWE-bench Pro, Aider-Polyglot.
4 benchmarks
Math Capability
Mathematical reasoning benchmarks: AIME 2025/2026, FrontierMath, MATH-500, GSM8K.
5 benchmarks
AI Agent Capability
Autonomous agent benchmarks: τ²-Bench, Terminal Bench Hard, Aider-Polyglot.
3 benchmarks
Reasoning Capability
Reasoning and thinking benchmarks: HLE, ARC-AGI-2, GPQA Diamond.
3 benchmarks
General Performance
General AI performance: MMLU-Pro, LMArena Elo ratings.
2 benchmarks
OpenClaw Ranking
OpenClaw agent performance: Claw Bench and Pinch Bench.
2 benchmarks
Comprehensive Ranking
Overall scores across HLE, ARC-AGI-2, FrontierMath, SWE-bench, and τ²-Bench