Anthropic

Claude Mythos

Introduction: A Sudden Leap Forward

On March 27, 2026, leaks began circulating about Anthropic testing a new large language model called Claude Mythos, described as a significant leap beyond the current flagship Claude Opus. What was initially rumored to be months away from release was officially announced by Anthropic hours later, alongside a cybersecurity initiative called Project Glasswing. However, this isn't a standard public launch—Anthropic explicitly stated that Mythos won't be available to the general public due to concerns over its powerful vulnerability discovery and exploitation capabilities.

The model had been under internal testing since February 24, 2026, but gained attention in late March after a configuration error in Anthropic's content management system accidentally exposed a draft blog post.

What is Claude Mythos?

From an architectural perspective, Mythos is a general-purpose language model, internally designated as "Capybara"—a new tier above Opus in Anthropic's lineup. It's larger, more capable, and more expensive than the current Claude Opus 4.6, with marked improvements in coding, reasoning, and autonomy.

Anthropic's official system card describes it as: "Substantially surpassing any model we have previously trained across diverse domains such as software engineering, reasoning, computer use, knowledge work, and research assistance."

Crucially, these capabilities aren't specifically trained for cybersecurity; instead, they are byproducts of general enhancements—the same improvements that make the model better at fixing vulnerabilities also make it better at exploiting them.

Benchmark Dominance: A Clear Step Above Opus 4.6

The official system card provides direct comparisons with Claude Opus 4.6. While Opus 4.6 already ranks among the top globally on most benchmarks, competing closely with models like GPT-5.4 and Gemini 3.1 Pro, Mythos operates at a distinctly higher level, leading significantly across evaluations.

Key metrics show Mythos achieving first place in nearly all benchmarks, including coding tasks where SWE-bench Pro scores jumped from 53.4% to 77.8%—a near 25-percentage-point increase. This benchmark is designed to minimize data contamination by using novel problems, making the gap harder to attribute to memorization.

Token Efficiency: Smarter, Not Just Brute Force

An interesting data point from Anthropic's technical report is the BrowseComp test, which assesses a model's ability to retrieve hard-to-find information online using tools like web search and code execution. Mythos Preview scored 86.9% versus Opus 4.6's 83.7%—a modest accuracy difference—but the real standout is token efficiency. Mythos used approximately 4.9 times fewer tokens to achieve similar performance, averaging 226,000 tokens per task compared to 1.11 million for Opus 4.6. This indicates more concise reasoning paths, leading to lower inference costs and faster responses, even though Mythos itself is priced higher.

Anthropic acknowledges potential data contamination in this benchmark, but argues that the model's improving curve with increased token budgets suggests memorization isn't the sole explanation.

Security Capabilities: Beyond Human Expertise

Mythos isn't just superior in coding; it demonstrates unprecedented autonomous security capabilities, surpassing human experts in some tasks. In red-team tests, Anthropic provided Mythos with isolated environments and software source code, prompting it simply with "find security vulnerabilities," then let it independently analyze, hypothesize, verify, and generate exploits.

Notable achievements include:

  • Under $50 cost: Discovered a 27-year-old zero-day vulnerability in OpenBSD involving TCP SACK implementation flaws, which could crash TCP-responding hosts.
  • Under $1,000 cost: Developed a full remote code execution exploit for a 17-year-old FreeBSD NFS vulnerability (CVE-2026-4747), autonomously completing the entire process from discovery to working exploit—something Opus 4.6 achieved only with human guidance.
  • Under $2,000 cost: Chained multiple Linux kernel vulnerabilities for privilege escalation, from bypassing KASLR to achieving root access, all without human intervention.
  • FFmpeg vulnerability: Found a 16-year-old H.264 out-of-bounds write bug that evaded years of fuzz testing and manual review.
  • Firefox exploit quantification: Against known JS engine vulnerabilities in Firefox 147, Mythos succeeded in creating working exploits 181 times, with 29 cases achieving register control, compared to Opus 4.6's 2 successes in hundreds of attempts.

This isn't merely assisting security researchers; it's automating top-tier red-team capabilities.

Why Mythos Isn't Public and Project Glasswing

Anthropic has no plans to release Mythos to the general public, citing direct weaponization risks. Logan Graham, Anthropic's former red-team lead, notes that Mythos Preview is "extremely autonomous" with advanced security researcher-level reasoning, enabling it to both find and exploit vulnerabilities—a step change from Opus 4.6, which had near-zero success in autonomous zero-day exploitation.

More importantly, Anthropic anticipates that other AI companies will release models with similar capabilities within 6 to 18 months. To address this, they've launched Project Glasswing, a defensive cybersecurity partnership with about 40 institutions, including core partners like AWS, Apple, Google, Microsoft, and NVIDIA. These organizations will use Mythos to scan their own and open-source software for vulnerabilities, with shared findings benefiting the broader industry. Anthropic is providing up to $100 million in API credits to participants and $4 million to open-source security groups like OpenSSF.

Mythos Preview is also available in private preview on Google Cloud Vertex AI for eligible customers within the Glasswing network.

Implications: A Capability Gap Revealed

A key takeaway is that Mythos is a general-purpose model; cybersecurity is just one area where its enhanced capabilities manifest. Anthropic's system card explicitly states that its vulnerability-hunting skills stem from improved code understanding, reasoning, and autonomy—not specialized training. Thus, the model that autonomously chains exploits overnight is the same one leading benchmarks like SWE-bench Pro with 77.8% and achieving 56.8% on Humanity's Last Exam.

This suggests Anthropic's actual capabilities likely far exceed what the public sees. Mythos began internal testing just two months after Opus 4.6 launched, yet it shows step-function improvements across nearly all dimensions. By choosing not to release it publicly, Anthropic signals that Mythos's abilities may outpace existing safety frameworks—a first for an AI company acknowledging in official documentation that they're grappling with controlling what they've built, not as a hypothetical risk, but as a current reality.

For the industry, Mythos compresses two time windows: the preparation period before AI-assisted attacks become widespread, and the time for competitors to catch up. Anthropic's red-team estimates this window at 6-18 months. Regardless of the timeline, Mythos has transformed AI security threats from a discussion point into a data-backed reality. The public still uses Opus 4.6, but Mythos's existence highlights a previously unknown gap between Anthropic's capability reserves and its released products.

Comments (0)

Share:XHatena

Post a Comment

Loading...