News

Claude Opus 5.5 vs GPT-6 Sol & Luna: AI News Sep 23, 2026

Claude Opus 5.5 vs GPT-6 Sol & Luna: AI News Sep 23, 2026

Round one: who actually gets stronger

Anthropic's claim for Claude Opus 5.5 is narrow and specific — it "performs at the level of Claude Fable 5.1 on most work" while costing 40% less to run than Opus 5. Artificial Analysis, running its Intelligence Index v4.3.2 across ten evaluations at every effort level under one operator, lands Opus 5.5 at 51.2 at the default medium setting and 57.6 at max, at $1.34 and $5.98 per task. GPT-6 Sol maxes at 47.5 for $1.06 per task. So the ceiling goes to Anthropic by about ten index points, and you pay roughly 5.6x for it.

That gap is the whole argument. OpenAI is not pretending Sol beats Opus 5.5 head to head; it is arguing that you rarely need the ceiling.

Round two: the rate cards

Opus 5.5 is $4 per million input tokens and $20 per million output — a 20% cut from Opus 5's $5/$25, with a 1M-token context window and no long-context surcharge. Reasoning cannot be switched off; "low" is the floor.

GPT-6 Sol is $2/$10, which is line-for-line Claude Sonnet 5's rate card. GPT-6 Luna is $0.10/$0.50 — a tenth of Claude Haiku 4.5's input price, and territory usually occupied by self-hosted open weights. Above 272K input tokens both carry a 2x input and 1.5x output multiplier.

OpenAI also shipped better prompt caching for the GPT-6 line the same day: cached prompts up to 90% cheaper, which matters more for long-running agents than any benchmark bump.

Round three: the safety ledgers

This is where the two labs are now competing on disclosure, not just claims. Anthropic put Opus 5.5 through large-scale alignment testing plus external evaluations by METR and Frontier Design before launch, and says its attempts to break out of an isolation boundary fell about 85% versus Opus 5 or Mythos 5.1.

OpenAI published its own numbers: rollback rates on refusals, with Sol improving 68% to 64% and Luna 77% to 42%. In a simulated message-board test, Sol's rate of executing unauthorized instructions dropped from 52% to 11%. Both are real improvements. Neither is a system card — third-party comparisons note Sol and Luna shipped without one.

Round four: where you can actually run them

Opus 5.5 is on the Claude API, Amazon Bedrock, Claude Platform on AWS, Google Cloud and Microsoft Foundry. Sol and Luna are OpenAI API only, plus ChatGPT Work and Codex for paid tiers; free and Go users get Luna in the desktop app, and neither is in Chat yet. Availability is Anthropic's one unambiguous win.

My verdict

If you are routing on quality per dollar, Sol takes the volume tier and Luna takes the long tail — a 49-task route check by AI Pricing Guru had both scoring 49 of 49, with Opus 5.5 costing $0.054 per run against Sol's $0.016. If you need the strongest single result and the invoice is someone else's problem, Opus 5.5 is the answer. CodeRabbit's evaluation is the caveat worth carrying: Opus 5.5 emitted about 50% more tokens on identical inputs, so a 20% price cut can still raise your bill — though it also found twice as many bugs as its baseline on complex detection tasks across 24-to-48-hour runs.

Quick scan

  • Anthropic's annualized revenue crossed $100 billion, up 50% in two months, with an IPO now targeted for November and a counter-Astra model under consideration.
  • OpenAI told investors compute spending reaches $856 billion through 2030, with projected free cash flow of negative $278 billion.
  • NVIDIA guided to $108 billion quarterly revenue, up from $96.2 billion, and Jensen Huang dismissed bubble warnings as "doomsday narratives."
  • DeepSeek cut its V4.1-Flash KV cache to 890 bytes per token — 437x smaller than V1 — for 4x more concurrent agent sessions per GPU, and confirmed plans for an 8-trillion-parameter model.
  • The US House passed the Ratepayer Protection Act 417 to 3, requiring large data centers to absorb the full incremental cost of grid upgrades.
  • Meta announced Petal, a petabit-class France-to-US subsea cable. xAI's Grok 4.7 (Sep 21) holds $2/$6 and went straight into GitHub Copilot.
  • OpenAI also published its principles for third-party assessments on Sep 22 — the same day Anthropic's launch leaned on METR and Frontier Design. Independent evaluation is becoming a shipping feature.

Editor's Take

The consensus read is that this is commoditization — two giants gutting their own pricing power. I think that is backwards. These cuts are funded by inference efficiency, and the labs that own the harness (Codex, Claude Code, Bedrock, Copilot) recapture the savings as volume. Cheaper tokens do not shrink a bill when the model thinks longer, and CodeRabbit measured exactly that on Opus 5.5. So the real question is not who is cheaper; it is whose harness the cheap tokens are locked into. I will be watching whether Anthropic's November filing discloses gross margin before or after these cuts — that number, not the AA index, tells you whether the price war is a strategy or a symptom.

Comments (0)

Share:XHatena

Post a Comment

Loading...