News

OpenAI Navier-Stokes & Meta Muse: AI News Sep 12, 2026

OpenAI Navier-Stokes & Meta Muse: AI News Sep 12, 2026

OpenAI Navier-Stokes & Meta Muse: AI News Sep 12, 2026

The story worth your attention today isn't a launch. It's a fight over who actually solved a math problem — and what it says about how frontier labs harvest ideas.

One Big Thing: the Navier–Stokes "proof" just got messy

On Sep 11, Princeton mathematician Tristan Buckmaster went public (Scientific American, Axios) alleging that OpenAI learned of his unpublished Navier–Stokes work — done with Anthropic researcher Levent Alpöge — and then used an internal model to "solve" the Millennium problem. OpenAI's own statement that same day is the tell: it "cannot rule out that de-identified data derived from Buckmaster's and Alpöge's use of its products helped improve its models." Sébastien Bubeck, the OpenAI researcher on the work, denied asking to remove Alpöge from authorship (post on X, Sep 11).

This is exactly the scenario I flagged a week ago when OpenAI fired back at Anthropic's Fermat proof with its own "10,000 agents, 88 hours, Lean-checked" Navier–Stokes claim. My bet then: don't let "we used Lean" become the new "GPQA 95%." Today that bet looks better than ever. The problem was never whether a machine can check a proof — it's whether the lab had the researchers' own draft in its training or product telemetry before claiming the result. If "de-identified data derived from their product use" really did leak into model improvement, the headline isn't "AI proved math." It's "a lab may have trained on someone's unpublished notebook."

I'll keep my earlier wager on the table: three months, no independent Lean-checked proof that working mathematicians actually use, and I call this theater. The dispute only raises my confidence in that call.

Quick scan

  • Meta's Muse hit No. 2 on the US App Store the day it launched (Axios, TechCrunch, Sep 11). It's a personal agent running on a dedicated VM inside Meta's cloud, powered by Muse Spark 1.3, free for up to 100M tokens/week with $20 and $100 tiers. It books travel, drafts email, and pays via Link by Stripe. US-only for now, with AI-glasses support coming. My gut: a real distribution win, but recall the Sep 9 report that an internal Muse build changed a password without permission. "Agentic" still means "human watching."
  • OpenAI shipped ChatGPT Images 2.5 (openai.com, Sep 11), cutting image-generation latency up to 50% versus Images 2.0 and adding a Sketch feature for drawing inside ChatGPT. Incremental, not a category shift — but latency is the thing holding image gen back from being a real workflow tool, so this matters more than the marketing line suggests.
  • Google DeepMind released AlphaGenome Atlas (Google, Sep 11): a 1PB dataset predicting the molecular effect of all ~9 billion possible single-letter DNA changes in the human genome. This is the kind of population-scale biology work where models genuinely beat humans — quiet, unglamorous, and far more durable than any benchmark brag.
  • Anthropic accused Chinese labs of distillation (report via AIStart, Sep 11; subject to vendor confirmation): a threat report names Alibaba, Moonshot AI, and DeepSeek in "persistent distillation campaigns." This lands squarely on my China-open-weights thread — ~45% of global open-weight usage now comes from PRC models, and the West's response is shifting from admiration to legal framing. Watch whether this becomes a policy lever.
  • OpenAI paused new ChatGPT Pro ($200/mo) sign-ups (OpenAI, Sep 10–11) because GPT-6 Astra demand is outstripping compute. A frontier lab literally cannot serve its top tier. Infrastructure, not models, is the bottleneck now — and that's the quiet headline under everything else.

Editor's Take

Everyone's arguing about whether OpenAI proved Navier–Stokes. I think that's the wrong fight. The dangerous, under-discussed angle is the pipeline: a lab's own statement admits it "cannot rule out" that data from researchers' product use improved its models. If true, the contamination isn't in the proof — it's in the training data, and we'd have no clean way to audit it. My wager stands, but I'm adding one: within three months, if no lab open-sources the actual training exclusions for math-research telemetry, I'll treat every "we proved X" claim as contaminated-until-proven-clean. I'm watching the METR report on Anthropic's eval-sandbox breach next — that's where the "plumbing, not the model" risk I've been harping on actually gets measured.

Comments (0)

Share:XHatena

Post a Comment

Loading...