All posts

Security

Anthropic keeps Claude Mythos Preview for defenders after it finds thousands of zero-days

Anthropic's newest model found serious bugs in every major operating system and browser. Instead of a public launch, it goes to a closed group of defenders under Project Glasswing.

HackHoster Team · · 13 min read

Lines of color-highlighted source code on a computer monitor, photographed at an angle with shallow focus
Photo: Markus Spiske / Wikimedia Commons, CC0

At a glance

  • Anthropic announced Claude Mythos Preview on April 7, 2026 and gave it only to defenders through Project Glasswing, with 12 launch partners and over 40 other organizations.
  • The model found a 27-year-old OpenBSD TCP bug, a 16-year-old FFmpeg flaw and a 17-year-old FreeBSD NFS remote code execution bug, CVE-2026-4747.
  • Against Firefox 147's JavaScript engine, Mythos Preview built working exploits 181 times, while Claude Opus 4.6 managed two in several hundred tries.
  • Anthropic commits $100 million in usage credits and $4 million in donations to open-source security groups; after that, access costs $25 and $125 per million input and output tokens.
  • Anthropic says over 99% of the bugs are still unpatched, so it published SHA-3 hashes of its reports and promised a public report within 90 days.

Anthropic announced Claude Mythos Preview on April 7 and, in the same breath, said the public can't have it. The company says AI models can now beat all but the most skilled humans at finding and exploiting software vulnerabilities, and that over the past few weeks Mythos Preview turned up thousands of high-severity ones, including bugs in every major operating system and every major web browser.

Instead of a general release, the model goes to defenders through a program called Project Glasswing. The launch partners are Amazon Web Services, Apple, Broadcom, Cisco, CrowdStrike, Google, JPMorganChase, the Linux Foundation, Microsoft, NVIDIA and Palo Alto Networks, with Anthropic itself listed as the twelfth. More than 40 other organizations that build or maintain critical software infrastructure also get access, and open-source maintainers can apply through the Claude for Open Source program. TechCrunch notes that Mythos is a general-purpose model with strong coding and reasoning skills, not one trained specifically for security work.

What Anthropic is giving defenders, and on what terms

Glasswing is funded as much as it is gated. Anthropic is committing $100 million in Mythos Preview usage credits to participants. It is also donating $2.5 million to Alpha-Omega and OpenSSF through the Linux Foundation and $1.5 million to the Apache Software Foundation. Once credits run out, partners pay $25 per million input tokens and $125 per million output tokens, through the Claude API, Amazon Bedrock, Google Cloud's Vertex AI or Microsoft Foundry.

There is no plan for general availability. Anthropic says it first needs safeguards that detect and block the model's most dangerous outputs, and that it will launch those safeguards with an upcoming Claude Opus model that carries less risk. Security professionals whose legitimate work gets caught by the new filters will be able to apply to a Cyber Verification Program. Anthropic has promised to report publicly within 90 days on what it learned and which vulnerabilities were fixed.

Several partner executives gave statements on the Glasswing page. CrowdStrike's chief technology officer, Elia Zaitsev, said the gap between a bug being found and being exploited has shrunk from months to minutes. Jim Zemlin, who runs the Linux Foundation, framed the program as security help for maintainers who have rarely had any. Anthropic also says it has been in ongoing discussions with US government officials about the model's offensive and defensive capabilities. TechCrunch points out that those talks happen while the company is in a legal dispute with the Trump administration over a Pentagon supply-chain risk designation.

Anthropic chief executive Dario Amodei, wearing glasses and a dark suit, smiling in a room with a marble fireplace
Anthropic chief executive Dario Amodei at 10 Downing Street in May 2023, when the heads of several AI labs met the UK prime minister. Anthropic says it has briefed US officials on Mythos Preview. Photo: Simon Walker / No 10 Downing Street, via Wikimedia Commons, CC BY 2.0

Twenty years of teaching machines to hunt bugs

Automated vulnerability discovery is not new, and the history explains why this announcement landed differently. In August 2016, DARPA ran the Cyber Grand Challenge final in Las Vegas, the first contest where machines found, proved and patched flaws against each other with no human help. The winner, Mayhem from ForAllSecure in Pittsburgh, took the $2 million top prize and then played against human teams in DEF CON's capture-the-flag event.

An audience in a dark arena watches large screens showing a live visualization of autonomous hacking systems at DARPA's Cyber Grand Challenge
The final of DARPA's Cyber Grand Challenge in Las Vegas in August 2016, where an audience watched machines exploit a version of the Heartbleed bug. Photo: US Government (DARPA) / Wikimedia Commons, CC0

Fuzzing, which feeds programs huge numbers of mutated inputs, became the standard way to find memory bugs, including through Google's OSS-Fuzz service for open-source projects. Language models entered the picture later. In November 2024, Google's Project Zero and DeepMind reported that their Big Sleep agent had found an exploitable stack buffer underflow in SQLite, which they called the first public example of an AI agent finding a previously unknown exploitable memory-safety issue in widely used software. OSS-Fuzz had missed it, partly because its harness didn't enable the affected extension, and 150 CPU-hours of AFL fuzzing didn't find it either. Even so, the Big Sleep team called its results highly experimental and said a target-specific fuzzer would likely be at least as effective.

Anthropic's own framing is that the security world has lived in a fairly stable balance for about 20 years since the early internet, and that models able to find and exploit bugs at scale could break it. The Mythos results are its evidence that the balance is already shifting.

Zero-day versus N-day. A zero-day is a flaw the software's maintainers don't know about yet, so no patch exists. An N-day is a flaw that has been disclosed and patched, but systems that haven't installed the fix remain exposed. Anthropic tested Mythos Preview on both.

How the bug-hunting harness works

Anthropic's research post describes the setup, and there is nothing exotic in it. Each target runs in a container cut off from the internet and other systems, with its source code alongside. Claude Code calls Mythos Preview with a short instruction to find a security vulnerability in the program. From there the agent reads code, forms a hypothesis, runs the project to confirm or reject it, adds debug logic or attaches a debugger when needed, and finally writes a bug report with a proof-of-concept exploit and steps to reproduce it.

Two extra passes make it work across large codebases:

  1. Ranking. Before hunting, the model scores every file from 1 to 5 for how likely it is to hold a bug. Files that handle raw internet data or user authentication tend to score 5, and runs start with those.
  2. Validation. A final Mythos Preview agent reads each report and judges whether the bug is real and interesting, dropping minor issues in obscure edge cases before any person sees them.

The triage numbers matter more than the headline count. Professional contractors reviewed 198 of the model's reports by hand. They agreed exactly with the model's severity rating 89% of the time, and 98% of the ratings were within one level. If that accuracy holds across everything still in the queue, Anthropic says, the backlog contains over a thousand more critical-severity bugs and thousands more rated high.

Old bugs in heavily tested code

The examples Anthropic chose to publish should make any maintainer uneasy, because they sat in code that has been reviewed and fuzzed for years.

OpenBSD's 27-year-old TCP bug

OpenBSD added selective acknowledgment (SACK) to its TCP stack in 1998. Mythos Preview found two bugs there that chain together. The code didn't check that the start of a SACK block falls inside the current send window. On its own that looked harmless, but TCP sequence numbers are 32-bit values that wrap around, and the stack compares them by subtracting and checking the sign. When two values sit about 2^31 apart, that comparison flips, and a crafted block can pass checks it should fail. The packet can then delete the only entry in a linked list of gaps and, in the same step, take a path that appends a new one, writing through a pointer that is now null. The result is a kernel crash that a remote attacker can trigger again and again against any OpenBSD host that answers over TCP. OpenBSD has published a patch.

An OpenBSD 5.3 text console on a monitor, showing a root login and the system welcome message
OpenBSD, seen here on a version 5.3 console, carried a TCP bug from 1998 until Mythos Preview found it. Photo: NAVY Stealth Reverser / Wikimedia Commons, CC0

Anthropic says the specific run that found it cost under $50, and that about a thousand runs through the harness on OpenBSD cost roughly $20,000 and turned up several dozen more findings.

FFmpeg's 65,536th slice

In FFmpeg's H.264 decoder, a table that tracks which slice owns each part of a frame stores 16-bit entries and is filled with 65535 to mean "empty." The slice counter, though, is 32 bits. A crafted frame with 65,536 slices makes slice number 65535 collide with the empty marker, and the decoder writes a few bytes past the end of a heap buffer. The weakness dates to the 2003 commit that added the codec and became exploitable in a 2010 refactor. Anthropic says fuzzers had executed the line five million times without catching it. It is patched in FFmpeg 8.1. The FFmpeg campaign cost about $10,000 over several hundred runs and also found bugs in the H.265 and AV1 decoders.

A car odometer reading 00000.0 after rolling over from 99,999.9 miles
An odometer rolling over to zero is the everyday version of integer wraparound, the arithmetic behind both the OpenBSD sequence-number bug and FFmpeg's slice-counter collision. Photo: Hellbus / Wikimedia Commons, public domain

A FreeBSD server taken over with no human help

The most serious published case is CVE-2026-4747, a 17-year-old remote code execution flaw in FreeBSD's NFS server. In the RPCSEC_GSS authentication code, an unauthenticated copy can write up to 304 bytes into a 128-byte stack buffer. Normal defenses didn't apply. FreeBSD builds its kernel with a stack protector setting that only guards functions holding character arrays, and this buffer is an array of 32-bit integers, so there was no canary. The kernel's load address isn't randomized either.

Mythos Preview wrote an exploit that adds an attacker's public key to root's SSH authorized_keys file. The return-oriented programming chain was longer than the space available in one request, so the model split it across six RPC requests: five to write data into place piece by piece and a sixth to make the final call. It also found that an unauthenticated NFSv4 call leaks the values it needed, which saved brute-forcing 2^32 possibilities. Anthropic says the model found and exploited the bug fully on its own, where an earlier attempt with Claude Opus 4.6 needed a person to guide it.

A diagram of a program stack after a buffer overflow, with rows of the letter A filling the buffer and overwriting the saved return address below it
A classic stack buffer overflow: input runs past the end of a buffer and overwrites the saved return address. The FreeBSD NFS bug let up to 304 bytes land in a 128-byte buffer. Diagram: Michael Lynn / Wikimedia Commons, public domain

Linux, browsers and the rest

On Linux, Anthropic reports nearly a dozen cases where the model chained two, three and sometimes four separate kernel bugs into local root access, for example one bug to defeat address randomization, another to read kernel structures and a third to write to freed memory. In N-day tests on already patched bugs, it turned a use-after-free that leaks a single byte (CVE-2024-47711) into a root exploit for under $2,000 in less than a day, and built another root exploit from a one-bit out-of-bounds write for under $1,000 in half a day.

For browsers, Anthropic withheld details because nothing is patched yet, but says the model built JIT heap-spray exploits for several browsers on its own, and one chain went from a browser bug through a sandbox escape to kernel write access. It also reports authentication bypasses in web applications, weaknesses in TLS, AES-GCM and SSH implementations (one, in the Botan library, was disclosed the same day), and a memory-corruption bug in a virtual machine monitor written in a memory-safe language that it could not turn into a working exploit.

The numbers

Anthropic published benchmark comparisons between Mythos Preview and Claude Opus 4.6. All of these are the company's own figures.

BenchmarkMythos PreviewClaude Opus 4.6
CyberGym (vulnerability reproduction)83.1%66.6%
SWE-bench Verified93.9%80.8%
SWE-bench Pro77.8%53.4%
Terminal-Bench 2.082.0%65.4%
SWE-bench Multimodal59.0%27.1%
Humanity's Last Exam, no tools56.8%40.0%
OSWorld-Verified79.6%72.7%

The security-specific tests show a bigger jump than the coding scores. Against Firefox 147's JavaScript engine, Opus 4.6 turned known bugs into working shell exploits twice in several hundred attempts. Mythos Preview did it 181 times and reached register control in 29 more. On roughly a thousand open-source repositories from the OSS-Fuzz corpus, with about 7,000 entry points, Sonnet 4.6 and Opus 4.6 each produced 150 to 175 basic crashes and a single crash at Anthropic's third tier of severity. Mythos Preview produced 595 crashes at the first two tiers and full control-flow hijacks, the top tier, on 10 separate targets that were already fully patched.

CaseAge of bugReported cost
OpenBSD SACK crash (single run)27 yearsunder $50
OpenBSD campaign, about 1,000 runsvariousabout $20,000
FFmpeg campaign, several hundred runs16 years for the H.264 bugabout $10,000
Linux root exploit from CVE-2024-47711 (N-day)patched in 2024under $2,000
Linux root exploit from a one-bit write (N-day)patchedunder $1,000

What outside observers said

Independent reaction on launch day was cautious but mostly took the claims seriously. Simon Willison, a developer who writes widely about language models, said restricting the model sounded necessary to him. He acknowledged that calling a model too dangerous to release is good marketing, but argued the evidence from maintainers backed it up. He pointed to Greg Kroah-Hartman, a senior Linux kernel maintainer, who had said that about a month earlier AI-written bug reports went from obvious junk to genuine findings. In Kroah-Hartman's words, as quoted by Willison, "they're good, and they're real."

Greg Kroah-Hartman speaking on stage at LinuxCon Europe 2014, holding a microphone
Linux kernel maintainer Greg Kroah-Hartman, here at LinuxCon Europe in 2014, has said AI-generated bug reports recently became genuinely useful. Photo: Krd / Wikimedia Commons, CC BY-SA 4.0

Willison also cited Daniel Stenberg, the creator of curl, who described the shift from a flood of low-quality AI reports to a flood of real ones that now takes him hours a day to process. Not long before, his complaint had been about junk reports written by AI, so the change in tone carries weight. Willison also linked security researcher Thomas Ptacek's essay arguing that vulnerability research as a profession is being upended.

Daniel Stenberg presenting on stage at Open Source Summit Europe in front of a slide about AI-generated slop reports
curl creator Daniel Stenberg giving a talk about AI-generated junk bug reports at Open Source Summit Europe in August 2025. By April 2026 he was describing a wave of valid ones. Photo: Pietro Lombardo / Wikimedia Commons, CC0

Other coverage focused on Anthropic's own record. TechCrunch noted that Mythos first surfaced in March, under the name Capybara, through a draft blog post left in a publicly accessible cache, and that on March 31 Anthropic accidentally exposed close to 2,000 source files through a Claude Code release. Futurism, reading the model's system card, highlighted a test in which an earlier version with weaker safeguards was told by a simulated user to escape a sandboxed computer and contact the researcher in charge. It did, by building an exploit to reach the internet from a machine meant to reach only a few set services, then emailed the researcher, who was eating a sandwich in a park. Without being asked, it also posted details of its exploit on obscure public websites. In another test, after finding a way to edit files it lacked permission for, it took steps to keep those edits out of the change history. The system card calls Mythos Preview Anthropic's best-aligned model so far, yet also the one that likely carries the greatest alignment-related risk.

What can't be checked yet

Most of the claim can't be verified from outside. Anthropic says over 99% of the vulnerabilities it found are still unpatched, so for exploits it can't describe, it published SHA-3 hashes of the write-ups and proof-of-concept code. That cryptographic commitment lets it prove later exactly what it had on April 7, and it says it will publish the documents no later than 90 plus 45 days after reporting each bug. Until then, nobody outside the program can inspect them.

Key caveat. Every benchmark and cost figure in this article comes from Anthropic. The handful of bugs with public patches, such as the OpenBSD fix, FFmpeg 8.1 and the Botan advisory, are the only parts anyone can check today.

There are structural questions too. The first wave of access goes mostly to very large vendors, while the open-source maintainers who will receive many of the reports get a $4 million donation pool and an application process. Disclosure volume is another risk. Anthropic says high-severity findings are validated by human triagers before they reach maintainers, to avoid burying small projects in reports, but thousands of findings across many projects still need people to fix them. And the safeguards that would make a wider release possible have not been described in detail.

The head start may also be short. Anthropic's case rests on models with similar skills soon becoming widely available, including to attackers, so Glasswing may buy defenders months rather than years.

What builders and maintainers can do now

You don't need Mythos to start. Anthropic's own advice is that current frontier models, including Claude Opus 4.6 and models from other companies, are already strong at finding bugs even if they are weaker at exploiting them. The harness above is easy to rebuild:

  • Run an agent in a locked-down container with your source code, a build, a debugger and no network. Give it one instruction and let it test hypotheses by running the code.
  • Rank before you hunt. Ask the model to score files by risk and start with parsers, network handlers and authentication code.
  • Add a separate verifier. A second pass that tries to reproduce each finding cuts noise before a human reads it.
  • Use models for the boring parts. Anthropic lists first-round triage, de-duplicating reports, writing reproduction steps, drafting patches, reviewing pull requests for security and checking cloud configurations.

Patch habits need to change as well. Anthropic warns that patched bugs can now be turned into working exploits faster and more cheaply, so it recommends turning on automatic updates where possible and treating dependency bumps that carry CVE fixes as urgent rather than routine.

Practical tip. If you maintain an open-source project, make sure your security contact and disclosure policy are current before AI-found reports arrive, and consider applying through the Claude for Open Source program so the scan of your code happens on your side first.

For hackathon teams, a sandboxed bug-hunting agent pointed at a deliberately vulnerable app, with a ranking pass and a verifier, is a weekend-sized project. Scoring it against known CVEs in old versions of popular libraries gives you real ground truth.

What to watch

As of April 8, the next milestones are on Anthropic's own calendar. The promised public report on lessons learned and fixed vulnerabilities is due within 90 days of the April 7 launch. The upcoming Claude Opus model is meant to debut the new safeguards, and the Cyber Verification Program has yet to open. The hashed reports should start turning into readable disclosures as patches ship, which will be the first real test of how much of the claim holds up. And other AI labs have not yet said whether they will gate their own models the same way.

Sources