News

Grok 4.7 & GPT-6 Sol Price War: AI News Sep 24, 2026

Grok 4.7 & GPT-6 Sol Price War: AI News Sep 24, 2026

The real story this week isn't a benchmark

Sunday xAI pushed Grok 4.7 out the door. Monday, within about ninety minutes of each other, Anthropic shipped Claude Opus 5.5 and OpenAI shipped GPT-6 Sol and Luna. Three frontier labs, two days, three flagships — and here's the part I actually find interesting: not one of them led with "we're the fastest." They led with safety.

Grok 4.7's launch post spends as much ink on its "best-calibrated safeguards to date" as on its coding scores. OpenAI's Sol/Luna materials lean on refusal-rollback rates and a simulated message-board test. Anthropic put Opus 5.5 through METR and Frontier Design before it shipped. When the three labs that are supposedly racing each other all show up to the fight talking about guardrails, the "race" framing starts to look like choreography.

I've said for weeks that access tiers and evaluation gates are commercial control valves dressed up as safety. This week is the cleanest evidence yet.

Grok 4.7: the newest contender

xAI's new model is a coding-and-knowledge-work specialist, not a general flagship play. The specs are concrete: a 500k-token context window, text-and-image in, text out, no output cap, and reasoning effort dials from low through xhigh. Pricing is $2 per million input, $0.50 cached input, $6 output below 200k prompt tokens, doubling above that. A "Fast" variant doubles output speed at double price.

On the numbers xAI published (vendor figures, treat as such): CursorBench 4.0 lands at 46.3% versus 40.4% for Grok 4.6 and 51.8% for Claude Fable 5.1 at max — so it closes the gap to Anthropic on coding without quite reaching it. DeepSWE v1.1 (high) rises to 71.0% from 65.2%. Terminal-Bench 4.0 jumps to 38.0% from 20.3% — still well behind Fable's 57.9%, but a real leap for a model that previously flopped at agentic tool use. EEBench 64.0% beats Fable's 56.4%.

The safety stack is the part worth watching. xAI claims a LatchBio biosafety score of 62.4% and says only 3.3% of high-risk dual-use prompts slip through on HackerBench v0.3. It also ships with native understanding of the Grok Bot harness — the agent runtime — which is the tell. Grok 4.7 isn't meant to sit in a chat box; it's meant to run as a teammate on a persistent cloud machine. Available now in Cursor, Grok Build, the Grok API, and third-party harnesses.

The price war is the sideshow

Everyone's staring at the rate cards — Sol at $2/$10, Luna at $0.10/$0.50, Opus 5.5 at $4/$20. I covered that yesterday. What's new is the coordination: three labs dropping flagships in a 48-hour window, each one pre-cleared by an outside evaluator. Independent assessment is becoming a shipping feature, not a regulatory afterthought. OpenAI even published its "principles for effective third-party assessments" on the same day Anthropic leaned on METR.

That's either a genuine safety culture shift or the most expensive PR move of the quarter. My money's on "both, and the PR is the point."

Quick scan

  • Alibaba refreshed the Qwen3.8-Max 0902 checkpoint: 2.4 trillion parameters, 1M-token context, CodeArena up 22 points to 1,691 at unchanged pricing; a Qwen3.8 27B also dropped. (Vendor figures.)
  • Meta's Muse Spark 1.3 reportedly trims token use about 25% and tool calls about 20% versus 1.2, alongside a Muse Voice Transcribe tool with speaker ID for 20-plus people. (Reported; vendor to confirm.)
  • A practitioner note claims OpenAI pulls the plug on the Sora API today, Sep 24. OpenAI's own news page shows no shutdown notice as of Sep 23 — so treat that as unconfirmed until the vendor says so.
  • Apple's iOS 27.2 beta adds a motion-data restriction for China that blunts the "shake-to-open-ad" redirect trick — a small but telling sign of how adversarial the ad-tech cat-and-mouse has gotten.

Editor's Take

The consensus is "frontier AI is commoditizing, prices are collapsing, competition is brutal." I think that read mistakes a pricing strategy for a market outcome. These labs aren't racing to the bottom; they're synchronizing at the top, and the new variable they compete on is who can show the most credible safety story without slowing the launch. My bet: within three months, at least one of these three will quietly walk back a safety claim once an independent audit publishes. The cadence this week wasn't competition — it was a cartel learning to smile for the regulators. I'll be watching the audit logs, not the benchmark charts.

Comments (0)

Share:XHatena

Post a Comment

Loading...