Back to Leaderboard

Reasoning Capability

Reasoning and thinking benchmarks: HLE, ARC-AGI-2, GPQA Diamond.

698 models

#ModelDeveloperOpen Source
1Claude Mythos PreviewAnthropic64.794.6Closed
2GPT-5.4 ProOpenAI58.783.394.4Closed
3Muse SparkMeta AI58.042.589.5Closed
4GPT-5.5 ProOpenAI57.284.6Closed
5Opus 4.7Anthropic54.775.894.2Closed
6Kimi K2.6Moonshot AI54.090.5Closed
7Qwen3.7-Max-Previewアリババ53.592.4Closed
8Claude Opus 4.6Anthropic53.066.391.3Closed
9GLM 5.1Zhipu AI52.3Closed
10GPT-5.5OpenAI52.285.093.6Closed
11GPT-5.4OpenAI52.177.192.8Closed
12Gemini 3.1 Pro PreviewGoogle DeepMind51.477.194.3Closed
13Kimi K2 ThinkingMoonshot AI51.0Closed
14Qwen 3.6 Plus Previewアリババ50.690.4Closed
15GLM-5Zhipu AI50.44.9Closed
16Kimi K2.5Moonshot AI50.211.8Closed
17Qwen3.6-Max-Previewアリババ50.290.4Closed
18GPT-5.2 ProOpenAI50.054.293.2Closed
19Qwen3-Max-Thinkingアリババ49.8Closed
20Claude Sonnet 4.6Anthropic49.058.389.9Closed
21Qwen3.5-27Bアリババ48.5Closed
22Gemini 3 Deep Think - 2620Google DeepMind48.484.6Closed
23Qwen3.5-397B-A17Bアリババ48.388.4Closed
24DeepSeek-V4-ProDeepSeek48.289.1Closed
25Gemini 3.0 Pro (Preview 11-2025)Google DeepMind45.845.191.0Closed
26GPT-5.2OpenAI45.554.292.4Closed
27DeepSeek-V4-FlashDeepSeek45.188.1Closed
28Grok 4 HeavyxAI44.4Closed
29Gemini 3.0 FlashGoogle DeepMind43.533.690.4Closed
30Opus 4.5Anthropic43.237.6Closed

About Benchmarks

HLE
総合知能テスト — 人間レベルの推論能力を測定
ARC-AGI-2
抽象的推論ベンチマーク — 新規パターンの汎化能力を測定
GPQA Diamond
Graduate-Level Google-Proof Q&A — 大学院レベルの科学的推論能力を測定