OpenAI

OpenAI Launches GPT-5.5 (Spud)

OpenAI officially released GPT-5.5, codenamed "Spud", on April 24th. Arriving just six weeks after GPT-5.4, this rapid release cycle underscores that leading AI labs are now following a rolling iteration model rather than holding back for massive version jumps. GPT-5.5 is immediately available to ChatGPT Plus, Pro, Business, and Enterprise users, as well as Codex users. GPT-5.5 Pro is available for Pro, Business, and Enterprise tiers. The API launch, however, has been postponed due to additional cybersecurity verification requirements, with OpenAI stating it will arrive "soon."

GPT-5.5 Key Specifications

GPT-5.5 supports a context window of up to 1 million input tokens, though the maximum output is still expected to be 128K. It's important to note that within Codex, GPT-5.5's maximum support is 400K; the 1M window is an experimental feature requiring manual activation.

Two versions have been released: the standard GPT-5.5, which offers different levels of "thinking" depth, and GPT-5.5 Pro, which accepts text and image input but only outputs text.

Within ChatGPT, GPT-5.5 is exposed as a "Thinking" mode with adjustable thinking duration. Plus and Business users can select between Standard and Extended modes, while Pro users gain access to Light and Heavy modes. Codex offers a separate Fast Mode, which reduces latency by 1.5x but costs 2.5x the standard price.

OpenAI co-founder Greg Brockman positioned GPT-5.5 as a "new class of intelligence," specifically emphasizing its ability to complete more complex, multi-step tasks with less external guidance. In practical terms, this means users can hand off a half-finished task and let the model decompose, plan, and execute it independently.

Officially, the key improvement areas are Agentic Coding, Computer Use, general knowledge work, and scientific research assistance.

Notably, OpenAI highlights that GPT-5.5's per-token latency is on par with GPT-5.4, but it consumes fewer tokens to complete the same task. This is a crucial detail given that GPT-5.5's API pricing has doubled compared to its predecessor.

Benchmarks: Leading in Agents, Lagging in Pure Reasoning

The benchmark data from OpenAI presents an interesting split.

In agentic knowledge work, GPT-5.5 scores strongly: 84.9% on GDPval (a benchmark covering 44 professional categories), 78.7% on OSWorld-Verified (autonomous computer control), and 98.0% on Tau2-bench Telecom (complex customer service workflows). It also claims top position among published models on the bioinformatics benchmark BixBench.

Compared to GPT-5.4, the largest improvements are on: ARC-AGI-2 (+11.7 points), MCP Atlas (+8.1 points), and Terminal-Bench 2.0 (+7.6 points). ARC-AGI-2 is a general reasoning benchmark designed to resist rapid saturation, so gains of this magnitude are significant.

However, in pure reasoning scenarios without tool use, the picture is less favorable. On Humanity's Last Exam, GPT-5.5 Pro scores 43.1%, which is below both Claude Opus 4.7 (46.9%) and Mythos Preview (56.8%). This indicates that while GPT-5.5 excels in agent execution and tool-use scenarios, OpenAI does not currently lead in pure academic reasoning independent of tools.

Third-party platform BenchLM.ai places GPT-5.5 5th out of 112 models with an overall score of 89/100. Its strongest area is Agentic tasks (2nd place), and its weakest is multimodal and grounded understanding (64th place), aligning with the above analysis.

Coding and Agents: Is the Double Price Justified?

The doubling of GPT-5.5's pricing is a major talking point. It's unclear if this is due to a larger model size or a strategic pricing decision. The new price now exceeds that of Opus 4.7 and makes it the most expensive flagship model currently available (with Claude Mythos being inaccessible).

For coding engineering, OpenAI claims GPT-5.5 better understands system architecture and fault nodes, knowing where to make changes and what the downstream impact will be. Early testing suggests tasks in Codex require fewer retries and consume fewer tokens. This claim requires verification; even if true, given the price hike, costs might not be lower unless token consumption is cut by half—a scenario that's met with skepticism.

For computer use, its 78.7% score on OSWorld-Verified is among the highest for publicly benchmarked models. Early test teams have used it for batch review of thousands of documents, and others have compressed weekly business reporting workflows by 5-10 hours.

Nvidia's internal testing describes GPT-5.5 as a "chief of staff" type tool—capable of driving AI agents that already function in employee roles. This isn't just talk; Nvidia VP Justin Boitano stated the model underwent weeks of internal testing.

API pricing is $5 per million input tokens and $30 per million output tokens, exactly double the rate of GPT-5.4 ($2.50/$15). However, when factoring in token efficiency gains, the effective cost increase will be lower than this multiplier. High-throughput use cases will need concrete testing to judge true cost-effectiveness.

Enhanced Scientific Research Capabilities

OpenAI specifically emphasized research assistance, with Chief Research Officer Mark Chen stating GPT-5.5 has "substantial improvements" in scientific and technical research workflows, citing drug discovery as a key application area.

A compelling case study involves an internally customized GPT-5.5 with a specialized reasoning framework contributing to a new proof for a Ramsey Number, a central object in combinatorial mathematics. This suggests the model's contribution goes beyond information synthesis to genuine participation in mathematical reasoning generation. However, this was a specialized internal version; its capabilities cannot be directly equated with the general product.

GPT-5.5 Launch Summary: API Pending, Web Accessible

OpenAI's official line is that API deployment "requires different safety safeguards," and the company is working with partners to meet the security needs of large-scale services. Hence, the API is not yet online.

In summary, GPT-5.5 is a targeted upgrade focused on agent execution capabilities, not a revolutionary leap across all fronts. The gains on ARC-AGI-2, MCP Atlas, Terminal-Bench 2.0, alongside OSWorld-Verified's 78.7%, point to a model that is more practical and usable for real-world engineering and task automation scenarios—a progression somewhat similar to the Qwen 3.6 upgrade. Perhaps, the general capabilities of large models have reached a plateau.

For more detailed specifications, you can refer to the GPT-5.5 and GPT-5.5 Pro model cards on DataLearnerAI.

Comments (0)

Share:XHatena

Post a Comment

Loading...