OpenAI's mid-tier GPT-6 model, released September 22, 2026 at $2/$10 per million tokens - line-for-line Claude Sonnet 5's rate card, and half of GPT-5.6 Sol. It runs a 1.05M-token context window with 128K max output. The published numbers sell cost, not a new ceiling: 33.2% on AutomationBench at $0.27 per task, 56.4% on Agents' Last Exam - while DeepSWE v1.1 at 68.8% sits below GPT-5.6 Sol's best.
Parameters
Undisclosed
Context Window
1.05M
License
Proprietary
Release Date
2026-09-22
Benchmark Performance
AA Intelligence Index
47.5
LMArena Elo
—
HLE
—
ARC-AGI-2
—
SWE-bench Verified
—
GPQA Diamond
—
MMLU-Pro
—
LiveCodeBench
—
AIME 2025
—
MATH-500
—
API Pricing
Input Price (per 1M tokens)
$2
Output Price (per 1M tokens)
$10
Billing Mode: standard
Strengths
- •Half of GPT-5.6 Sol's price at $2/$10, with context raised to 1.05M tokens
- •33.2% on AutomationBench at $0.27 per task - the launch's most honest cost disclosure
- •Cached input at $0.20 per million (90% off) cuts the effective rate for long-running agents
Weaknesses
- •On DeepSWE v1.1 (68.8%) and OSWorld 2.0 it misses GPT-5.6 Sol's peak - what improved is cost, not the ceiling
- •All figures are vendor-reported at xhigh/max effort, and no system card shipped with the launch
- •Above 272K input tokens a 2x input / 1.5x output multiplier applies
Use Cases
- •High-volume agentic workflows with well-scoped tasks
- •Code review, test fixes and small features inside Codex
- •Cost-sensitive multi-step tool calling at scale
Deep Analysis
Artificial Analysis Intelligence Index
39.8 (medium) / 47.5 (max)
AA Index v4.3.2; Claude Opus 5.5 reaches 57.6 at max effort
Measured Cost per Task
$0.25 (medium) / $1.06 (max)
Roughly one fifth of Opus 5.5's cost at matched medium effort
Input / Output Price
$2 / $10 per 1M tokens
50% below GPT-5.6 Sol's promotional rate; cached read $0.20, cache write $2.50
Context Window
1.05M tokens
128K max output; above 272K input, rates double and output goes to 1.5x
Safety Movement
Refusal rollback 68% -> 64%
Unauthorized-instruction execution in a simulated message-board test fell 52% -> 11%
Release
Sep 22, 2026
Knowledge cutoff Apr 20, 2026; OpenAI API, ChatGPT Work and Codex at launch
Strengths
- ・Rates match Claude Sonnet 5 line for line while OpenAI claims roughly half the error rate of its predecessor
- ・On OpenAI's AutomationBench it scored 6.3% above Claude Opus 5 at about 9% of the per-task cost
- ・Refusal rollback and unauthorized-instruction numbers improved materially, and the 52% to 11% result is a large jump
- ・Prompt caching for the GPT-6 line improved the same day, cutting cached-prompt cost by up to 90% for long-running agents
Weaknesses
- ・About ten Artificial Analysis index points below Opus 5.5 at matched max effort, so the ceiling is genuinely lower
- ・No system card published at launch
- ・Runs only on the OpenAI API plus ChatGPT Work and Codex; not available in ChatGPT itself yet
- ・A 272K-token input threshold doubles input and cache rates and raises output 1.5x, which Anthropic does not charge
Competitor Comparison
| Model | Arena | SWE | GPQA | Price |
|---|---|---|---|---|
| Claude Opus 5.5 | AA 57.6 (max) | N/A | N/A | $4/$20 |
| GPT-5.6 Sol (promo) | AA 59 | N/A | N/A | $4/$20 |
| Claude Sonnet 5 | N/A | N/A | N/A | $2/$10 |
GPT-6 Sol is the workhorse half of OpenAI's September 22 pair, and its argument is arithmetic rather than benchmarks. At $2 per million input tokens and $10 per million output it lands exactly on Claude Sonnet 5's rate card, 50% below GPT-5.6 Sol's promotional pricing, and it inherits training methods from the GPT-6 Astra flagship. Artificial Analysis measures it at 39.8 on the Intelligence Index at default medium effort ($0.25 per task) and 47.5 at max ($1.06 per task) — roughly ten points under Claude Opus 5.5, at about a fifth of the cost. OpenAI also published concrete safety movement: refusals rolled back 68% to 64%, and unauthorized-instruction execution in a simulated message-board test dropped from 52% to 11%.
Sources
Analysis generated: 2026-09-23