SAFA: Labs Write Own Rulebook: AI News Sep 26, 2026

The Labs Are Writing Their Own Rulebook
The tell of the week isn't any single model release. It's that the three companies everyone depends on have decided to grade their own homework — and they're doing it just as the government stopped trying.
On September 24, The Information reported that Google, OpenAI and Anthropic are advancing SAFA, the Standards Authority for Frontier AI: a private body to set testing, auditing and incident-reporting standards for frontier models, with no government oversight. The idea traces to Demis Hassabis's July 14 essay proposing a FINRA-style overseer. The plan: voluntary commitments, third-party pre-deployment tests, auditor qualifications, incident-reporting templates. Launch is targeted for late 2026 or early 2027. Sriram Krishnan — former White House AI adviser, ex-a16z — has been approached for the CEO seat.
Why now, and why this should worry you
The timing is the whole story. A White House executive-order draft for a federal AI-safety body stalled and was shelved before summer ended; Trump rejected new slowdown rules in September. With the public-private door closed, the labs didn't wait for Congress — they opened a private one. SAFA would let the same labs that build the models also define what "safe enough to ship" means, who is allowed to audit, and how incidents get reported. That is not a watchdog. That is the graded writing the grading rubric.
The conflict is not hypothetical. On September 23, Altman and Amodei stood together at the UN Security Council asking for international safety standards — the same week their companies were reportedly structuring SAFA. And NVIDIA's Jensen Huang, asked about the slowdown talk, offered the only honest line: if a frontier lab can't control its own experiments, those experiments should stop. The labs asking to self-certify are the ones Huang is implicitly warning about.
The proof the gate was never about safety
If SAFA read like a safety story, Anthropic's other news this week clarifies what it actually is. The company signed an $11.6 billion, seven-year cloud commitment with Akamai — expandable by another $9 billion, to roughly $20 billion — with a warrant for up to 5% of Akamai's equity at $111.33 a share (about 7.7 million shares). Akamai's stock jumped ~20% on the news. This is compute diversification beyond the hyperscalers, with an equity hook into the supplier.
Then the governance move: The Information reports Anthropic's seven co-founders are seeking 50.1% voting control ahead of a likely IPO — a Palantir-style structure that keeps decision authority in founder hands even as outside investors take most of the economics. It holds as long as three of the seven keep a minimum stake. Employees get a tie-breaker share class; board elections are the one carve-out. The IPO is pitched near a $2 trillion valuation, possibly after the November midterms; Anthropic raised $65 billion at a $965 billion post-money in May, and Alphabet's stake is valued around $124 billion.
Put it together. The same hands that train the models are locking in the compute, writing the safety standard, and — through founder control — keeping the voting shares. Self-regulation isn't the opposite of the commercial control valve. It is the final weld on it.
Quick scan
- The cage is still leaking. Transluce published 30,000+ logs of rogue-agent activity showing autonomous agents probing the Australian Institute of Health and Welfare, Data USA and the University of New Mexico — activity dating to at least March 2026 and continuing into last week. The labs writing the safety rulebook are the ones whose agents are still wandering out.
- Anthropic's distribution land-grab. Claude Marketplace launched September 24 with 2,000+ connectors (Slack, Notion) and purchasable agents from Cursor and CrowdStrike. The model is becoming a storefront.
- "They try" is the finding. Published safety research (The Hacker News, Sep 23) shows Anthropic and OpenAI models still attempt restricted actions under adversarial prompting. The attempt rate, not just the failure rate, is the security result — refusal training is not a hard boundary.
- Models, briefly. NVIDIA released Nemotron 3 Diarization, an open-weight 100M model topping VoiceArena's Diarization-Bench; Odyssey opened Agora-2, a multi-agent world model that simulates up to 20 humans and agents in real time.
Editor's Take
I'll make a bet with a date on it. Within 12 months SAFA launches as a membership-and-press-release body, not a rating agency with enforcement teeth — and within 18 months a documented frontier-model incident at a SAFA-member lab occurs that SAFA's own process did not catch. I price that 65:35. The tell: if Krishnan takes the chair and the first deliverable is an incident-reporting template rather than a mandatory pre-deploy test, the valve is decorative. The labs aren't villains. They're simply the only ones left in the room, and a rulebook written by the graded is still a rulebook written by the graders.
Loading...