OpenAIProprietary

GPT-5.5

Compare this model

A high-performance foundation model developed by OpenAI. It features advanced reasoning capabilities and broad task compatibility. It is the predecessor to GPT-5.6 Terra.

Parameters

Undisclosed

Context Window

License

Proprietary

Release Date

2026-04-23

Japanese Language Capability

High-Quality JP

Multilingual model with strong Japanese language processing capabilities.

API Pricing

Input Price (per 1M tokens)

$2.5

Output Price (per 1M tokens)

$

Billing Mode: standard

Strengths

    Weaknesses

      Use Cases

        Deep Analysis

        Arena Elo

        1474

        #3 verified on BenchLM; Coding sub-arena: 1507

        SWE-bench Verified

        88.7%

        vs Claude Opus 4.7: ~87.6%

        Artificial Analysis Index

        60

        #1 at release; breaks three-way tie with Anthropic and Google

        Terminal-Bench 2.0

        82.7%

        vs Claude Opus 4.7: 69.4%; state-of-the-art

        GPQA Diamond

        93.6%

        vs Claude Opus 4.7: 94.2%, Gemini 3.1 Pro: 94.3%

        Input / Output Price

        $5 / $30 per 1M

        2x GPT-5.4; Pro tier: $30 / $180

        Strengths

        • Industry-leading agentic coding (82.7% Terminal-Bench 2.0, 88.7% SWE-bench Verified) with best-in-class tool-call reliability (97.4% first-attempt success)
        • Massive long-context improvement — 74.0% at 1M tokens on MRCR v2, up from 36.6% on GPT-5.4 — with a genuine 1M-token API context window
        • Most mature enterprise ecosystem: first-party SDKs (Python/TS/Java/Go/.NET), Azure OpenAI integration, SOC 2/HIPAA/GDPR compliance, and dominant ChatGPT distribution

        Weaknesses

        • Highest frontier API pricing at $30/1M output tokens; hidden reasoning-token inflation can multiply effective cost of high-effort calls by 3-5x
        • 86% hallucination rate on AA-Omniscience (vs Claude Opus 4.7's 36%) — weakest factual reliability among frontier models on broad knowledge tasks
        • Loses to Claude Opus 4.7 on SWE-bench Pro (58.6% vs 64.3%) and on multi-file refactor tasks; trails on conversational preference (LMArena Elo ~1474 vs ~1492)

        Competitor Comparison

        ModelArenaGPQAPrice
        Claude Opus 4.7~149294.2%$5/$25 per 1M
        Gemini 3.1 Pro~147094.3%$2–3.50/$12 per 1M
        DeepSeek V4 Pro~1450N/A~$0.30/$1.10 per 1M

        GPT-5.5, released April 23, 2026, is OpenAI's most capable general-purpose model and the first fully retrained base since GPT-4.5. Built on a natively omnimodal architecture and co-designed with NVIDIA's GB200/GB300 NVL72 hardware, it matches GPT-5.4's per-token latency despite substantially higher intelligence. The model leads the Artificial Analysis Intelligence Index (score: 60), achieving state-of-the-art results on agentic coding benchmarks like Terminal-Bench 2.0 (82.7%) and setting a new OpenAI mark on SWE-bench Verified (88.7%). Its configurable reasoning effort dial — spanning none through xhigh — lets developers trade latency and cost for depth on a per-call basis.

        The model's positioning centers on agentic work: multi-step coding, computer use, long-horizon research, and tool-calling agent loops. Codex, OpenAI's coding agent, is the primary deployment surface, with over 85% of OpenAI's own employees using it weekly. GPT-5.5 demonstrates a genuine leap in long-context retention (74.0% at 1M tokens, up from 36.6% on GPT-5.4) and ARC-AGI-2 abstract reasoning (85.0%, up from 73.3%). The architecture reportedly uses mixture-of-experts with community estimates of 100–200B active parameters, though OpenAI has not disclosed specifics.

        GPT-5.5 is priced at a premium — $5/$30 per 1M input/output tokens, double GPT-5.4's rates — though OpenAI argues that ~40% token efficiency gains partially offset the increase. The Pro variant ($30/$180) targets the hardest reasoning tasks as an escalation tier. The model's main competitive weaknesses are its high hallucination rate on broad knowledge tasks relative to Claude Opus 4.7, its loss to Opus on SWE-bench Pro (the harder, less-memorized coding benchmark), and the cost premium that makes naive migration expensive. For teams building agentic systems, tool-using agents, or computer-use workflows, GPT-5.5 is currently the strongest available foundation model from OpenAI.

        Analysis generated: 2026-07-17