Back to LeaderboardHLE 総合知能テスト — 人間レベルの推論能力を測定 ARC-AGI-2 抽象的推論ベンチマーク — 新規パターンの汎化能力を測定 FrontierMath - Tier 4 高度な数学問題 — 研究レベルの数学的推論能力を測定 SWE-bench Verified 実践的ソフトウェア開発タスク — 実際のバグ修正能力を測定 τ²-Bench 自律エージェントタスク — ツール呼び出しと推論の組み合わせ能力を測定
Comprehensive RankingCoding CapabilityMath CapabilityAI Agent CapabilityReasoning CapabilityGeneral PerformanceOpenClaw Ranking
Comprehensive Ranking
Overall AI model ranking across HLE, ARC-AGI-2, FrontierMath, SWE-bench Verified, and τ²-Bench.
…