xAIProprietary

Grok 4

Compare this model

Grok 4 is an inference model released by xAI on July 10, 2025. It accepts text and image inputs and generates text outputs. It features a 256K token context window, with core capabilities in reasoning and multilingual support.

Parameters

Undisclosed

Context Window

License

Proprietary

Release Date

2025-07-10

API Pricing

Input Price (per 1M tokens)

$3

Output Price (per 1M tokens)

$

Billing Mode: standard

Strengths

  • Multilingual support
  • Large 256K context window
  • Advanced reasoning capabilities

Weaknesses

  • Closed-source license limits flexibility
  • Output limited to text only
  • New model with an unestablished track record

Use Cases

  • Reasoning and analysis on multilingual text
  • Analyzing image data
  • Developing API-integrated applications

Deep Analysis

Intelligence Index

73

Artificial Analysis index; frontier-tier reasoning and coding

GPQA Diamond

87.5%

Graduate-level science Q&A; among the strongest launch results

SWE-bench Verified

72.5%

Resolving real GitHub issues with human verification

AIME 2025

91.7%

Competition-level mathematics

Input Price

$3/1M

Output $15/1M; prompt-cache $0.75/1M (75% off); 2x above 128K tokens

Context Window

256K tokens

Text + image input; knowledge cutoff Dec 31, 2024

Strengths

  • Frontier reasoning: 87.5% GPQA Diamond and 91.7% AIME 2025 place it among the strongest models released in mid-2025.
  • Cost-efficient for a frontier model: $3 input / $15 output per million tokens is far below Claude Opus 4 ($15/$75), and prompt caching cuts repeated prefixes to $0.75/1M (75% off).
  • Accepts image input with real-time X (Grok) integration and strong multilingual support, useful for research and live-data tasks.

Weaknesses

  • 256K context is smaller than Gemini 2.5 Pro's 1M window, limiting very-long-document, video-transcript, or large-codebase work.
  • Throughput (~44-75 tokens/s) and latency trail OpenAI o3 and Claude, making it less suited to interactive, latency-critical chat.
  • No open weights; the $300/month SuperGrok Heavy tier is the most expensive mainstream consumer subscription, and the knowledge cutoff is Dec 2024.

Competitor Comparison

ModelArenaSWEGPQAPrice
Claude Opus 4 (Anthropic)N/ASWE-bench Verified 72.5%79.6%$15/$75
Gemini 2.5 Pro (Google)N/ASWE-bench Verified ~72%86.5%$1.25/$10
GPT-5 (OpenAI)N/ASWE-bench Verified ~72%~85%$1.25/$10

Grok 4 is xAI's flagship reasoning model, released July 9-10, 2025. It pairs a 256K-token context window (text plus image input) with frontier-tier benchmark scores - 87.5% GPQA Diamond, 91.7% AIME 2025, and 72.5% SWE-bench Verified - at a comparatively low $3/$15 per million tokens. xAI markets it as a generalist reasoning engine with real-time access to the X platform, strong multilingual coverage, and a multi-agent 'Heavy' mode for the hardest queries. It undercuts Claude Opus 4 on both price and most reasoning benchmarks, though its context length and throughput lag Google's and OpenAI's top tiers.

Analysis generated: 2026-08-24