xAI's coding-and-knowledge-work flagship, released September 21, 2026. It keeps Grok 4.6's $2/$6 per-million pricing while extending reinforcement learning onto harder multi-hour tasks, improving self-verification, long-context handling and native understanding of the Grok Bot agent harness. The launch led with safety: a new safeguard stack that xAI says tops LatchBio biosafety at 62.4% and passes only 3.3% of high-risk dual-use prompts on HackerBench v0.3. On benchmarks it closes the gap to Claude Fable 5.1 on coding without reaching it - CursorBench 4.0 46.3% vs 51.8%, Terminal-Bench 4.0 38.0% vs 57.9% - but posts a real agentic leap over its predecessor, nearly doubling Terminal-Bench.
Parameters
Undisclosed
Context Window
500K
License
Proprietary
Release Date
2026-09-21
API Pricing
Input Price (per 1M tokens)
$2
Output Price (per 1M tokens)
$6
Billing Mode: standard
Strengths
- •Terminal-Bench 4.0 more than doubles from 20.3% to 38.0%, covering the area where Grok 4.6 was weakest
- •46.3% on CursorBench 4.0 at $2/$6, closing most of the gap to Fable 5.1's $10/$50 while costing a fifth
- •Best-calibrated safeguards to date per xAI: LatchBio biosafety 62.4%, only 3.3% of high-risk dual-use prompts slip through
Weaknesses
- •Still trails Claude on headline coding and office benches - Fable 5.1 tops CursorBench at 51.8% and Terminal-Bench at 57.9%
- •Same sticker as Grok 4.6, so no per-token efficiency gain; the 2x Fast variant is the only speed-up, at double cost
- •Native Grok Bot harness understanding ties the model to xAI's agent runtime rather than open tool ecosystems
Use Cases
- •Long autonomous coding sessions where Terminal-Bench gains matter
- •High-volume code review and debugging at the CursorBench tier
- •Regulated or security-sensitive assists needing biosafety and dual-use gating
Deep Analysis
Context Window
500K tokens
Text and image input, text output, no output-length cap
API Price
$2 / $6 per 1M
Input/cached-input/output below 200K tokens; $4/$12 above; Fast variant doubles both
CursorBench 4.0
46.3%
vs Grok 4.6 40.4%, Claude Fable 5.1 (max) 51.8% (xAI-reported)
Terminal-Bench 4.0
38.0%
up from Grok 4.6's 20.3%; Claude Fable 5.1 57.9% (xAI-reported)
EEBench
64.0%
beats Claude Fable 5.1's 56.4% (xAI-reported)
LatchBio biosafety
62.4%
New safeguard stack; only 3.3% of high-risk dual-use prompts pass on HackerBench v0.3
Strengths
- ・Agentic leap: Terminal-Bench 4.0 more than doubles from 20.3% to 38.0%, the area where Grok 4.6 was weakest.
- ・Coding at a low price: 46.3% on CursorBench 4.0 at $2/$6 undercuts Claude Fable 5.1's $10/$50 while closing most of the gap.
- ・Best-calibrated safeguards to date per xAI, with a LatchBio biosafety score of 62.4% and only 3.3% of risky dual-use prompts slipping through.
Weaknesses
- ・Still trails Claude on the headline coding and office benches: Fable 5.1 maxes CursorBench at 51.8% and Terminal-Bench at 57.9%.
- ・Same sticker as Grok 4.6, so no efficiency gain per token; the 2x Fast variant is the only way to go faster, at double cost.
- ・Native Grok Bot harness understanding ties the model to xAI's agent runtime rather than open tool ecosystems.
Competitor Comparison
| Model | Arena | SWE | GPQA | Price |
|---|---|---|---|---|
| Claude Fable 5.1 | — | 71.0% (DeepSWE v1.1 high) | — | $10/$50 per 1M |
| GPT-5.6 Sol | — | — | 94.6% (GPQA Diamond) | $5/$30 per 1M |
| Grok 4.6 | — | 65.2% (DeepSWE v1.1 high) | — | $2/$6 per 1M |
Grok 4.7 is xAI's coding-and-knowledge-work flagship, released September 21, 2026. It holds Grok 4.6's $2/$6 pricing while extending reinforcement learning onto harder, multi-hour tasks, improving self-verification, long-context handling, and native understanding of the Grok Bot agent harness. The launch led with safety: xAI says its new safeguard stack tops LatchBio biosafety at 62.4% and passes only 3.3% of high-risk dual-use prompts on HackerBench v0.3. On benchmarks it closes the gap to Claude Fable 5.1 on coding without reaching it — CursorBench 4.0 46.3% vs 51.8%, Terminal-Bench 4.0 38.0% vs 57.9% — but posts a real agentic leap over its own predecessor. It ships in Cursor, Grok Build, the Grok API, and third-party harnesses, with a 500K-token context and a 2x-speed Fast variant at double price.
Sources
Analysis generated: 2026-09-24