Anthropic launches OSS Scanner: free Claude Mythos security scans for open source, with no human review
Anthropic's new OSS Scanner gives critical open source projects free, periodic vulnerability scans from its strongest models. Reports skip human review. Here is what the numbers say.

Anthropic wants to do for language models what Google's OSS Fuzz did for fuzzing. On 8 October 2026 the company launched OSS Scanner, an opt in service that gives critical open source projects regular vulnerability scans from its strongest models, including Claude Mythos, at no cost. The twist is in the fine print: every report the scanner sends is fully model generated, with no human review or triage before it reaches a maintainer.
That one design choice is the whole story. For years open source maintainers complained that AI tools flooded them with confident nonsense. Now Anthropic is betting that its models are good enough to skip the human checkpoint entirely, and that speed matters more than polish when attackers can turn a bug into an exploit in minutes. This article walks through what the service does, the numbers Anthropic published, what early maintainers said, and the open questions that remain.
The primary source is Anthropic's own announcement, Launching an opt in vulnerability finding service for open source software. Independent coverage came from The Verge, which highlighted the trade off between faster alerts and the lack of human review.
Why now: models got good at finding bugs
Anthropic opens with a benchmark. On CyberGym, an academic benchmark for finding vulnerabilities, language models went from finding under 20 percent of the bugs at the start of 2025 to finding over 85 percent this year. That is not a gradual climb. It is the kind of jump that changes who has the advantage in software security.

According to Anthropic, this shift is already visible in maintainers' inboxes. Open source projects went from receiving mostly low quality AI output to receiving bug reports that are genuinely useful. That matches what we have seen elsewhere. Earlier this month Google froze parts of its open source bug bounty because of a wave of weak AI submissions, a story we covered in Google pauses its OSS VRP over AI slop. Both things are true at once: low effort AI reports are a burden, and high quality AI reports are becoming a real defensive tool. OSS Scanner is Anthropic's attempt to land firmly on the second side.
The bottleneck is people, not models
The most revealing paragraph in the announcement is about scale. Over the past six months, Anthropic used its latest models to scan some of the world's most important software. The result: more than 29,000 candidate vulnerabilities. Of those, Anthropic's team has only managed to manually review and triage about 6,000.

Six thousand verified bugs is a lot. But it also means roughly four out of five findings were sitting in a queue waiting for a human. Anthropic says maintainers increasingly asked for everything at once. After receiving the first verified reports, many asked for a bulk delivery of all unverified findings plus proposed patches. To date, Anthropic says it has sent nearly 5,000 reports this way, at the maintainers' own request, even though they had not been validated.
That behaviour is the real origin of OSS Scanner. Maintainers decided that a fast, possibly imperfect report beats a perfect report that arrives months later. Anthropic is now turning that informal fast track into a formal service. The existing coordinated vulnerability disclosure process, where humans verify findings before sending them, stays in place, especially for projects that lack the people to triage reports themselves.
What a report contains
Each OSS Scanner report is designed to be actionable without a long back and forth. According to Anthropic, a report includes:
- A self contained reproducer, so a maintainer can trigger the bug directly
- An explanation of the vulnerability and why it matters
- A bisection, where possible, showing when the bug was introduced
- A candidate patch, when one is available

The reproducer is the important part. A maintainer does not have to trust the model's reasoning. They can run the test and see whether the bug is real. One of the testers, OpenSSL Corporation's Anton Arapov, made exactly this point: when a report comes with a working exploit attached, verifying it is basically the whole job.
Scans are periodic, not one off. Projects that join get repeated audits as their code changes, which is closer to how OSS Fuzz works than to a single security review.
How accurate are unreviewed reports?
This is the obvious question, and Anthropic answers it with its own validation data. Before the launch, it asked the expert penetration testers who normally review its disclosure findings to check 97 critical and high severity vulnerabilities produced by the scanner across 48 projects.

The breakdown:
- 85 of 97 (88 percent) met the bar for Anthropic's formal disclosure process.
- 11 were real issues but duplicated known bugs or other findings from the same scan.
- 1 was invalid, a true false positive.
In other words, only about one percent of the high and critical findings in that sample were outright wrong. Anthropic adds that maintainers who received findings have seldom said a high or critical report was invalid. It also admits the weaknesses honestly: some maintainers said severity ratings can be inflated, or that the scanner misunderstood a project's threat model. Anthropic says it cannot guarantee perfection and will keep refining the system based on feedback.
It is worth stressing that these are Anthropic's own numbers. They were checked by Anthropic's own reviewers on a sample Anthropic chose. They are encouraging, but they are not an independent audit.
What maintainers said
Anthropic spent several weeks validating the pipeline with dozens of open source projects. Those first disclosures contained hundreds of bug reports, including several vulnerabilities that could be chained into unauthenticated remote code execution. That is the most serious class of bug: an attacker who needs no login can run their own code on the target.

The quotes Anthropic published are notably positive:
- wolfSSL (Todd Ouska): of 74 reports received, all but two were valid, and five became CVEs. With patches attached, the reports slotted straight into the existing fix process.
- PostgreSQL (Noah Misch): an unusually high share of findings uncovered real defects, several patches were usable nearly as is, and fast track access let the team fix new issues before they reached a stable release.
- OpenSSL Corporation (Anton Arapov): early AI reports about 18 months ago were appalling, but Anthropic's reports, raw model output included, were as good as and sometimes better than reports from humans.
- HotCRP (Eddie Kohler): the reports were thorough and clear, and showed a strong understanding of a complex permission model.
These are hand picked testimonials, of course. Still, getting maintainers of PostgreSQL, OpenSSL and wolfSSL on record matters. These are exactly the projects that sit under huge parts of the internet, and their maintainers are not known for tolerating noise.
Who can join
OSS Scanner is not open to every repository. Eligibility follows criteria similar to OSS Fuzz: a project should have a critical impact on infrastructure and user security. Core maintainers enroll by submitting a pull request to Anthropic's GitHub repository using a standard project template, and Anthropic decides case by case. There is an extended FAQ for details.
Anthropic also points to two related programs. A new Cyber Verification Program gives qualifying security professionals access to advanced cyber capabilities with fewer blocking classifiers. And Claude for OSS offers free Claude Max 20x subscriptions to help maintainers fix vulnerabilities and improve their projects. On the commercial side, Claude Security remains the paid product aimed at enterprises.
The bigger picture: an arms race on a clock
Anthropic frames the service as defensive. Its argument is simple. Exploits can now be developed in minutes. Attackers are racing to find the same weaknesses with the same kind of tools. The projects that learn about a bug first and patch it fastest will be the ones that stay safe. Waiting for a human to verify every finding means giving attackers a head start.
The context backs that up. In May, AI tools helped surface the "Copy Fail" bug that affected nearly every Linux distribution, as The Verge notes. Just this week, CrowdStrike described a likely single attacker who used AI agents to breach several South Korean banks. The same capabilities that let Anthropic find 29,000 candidate bugs are available, in weaker or stronger form, to people with worse intentions.
There is also a strategic angle. Running the strongest models for free on critical code is expensive, and it builds goodwill with the developers who decide which tools their communities use. It mirrors what Google gained from OSS Fuzz: a reputation as a steward of open source security.
Open questions
Several things are still unclear:
- Maintainer workload. Even accurate reports take time to fix. A small project that suddenly receives dozens of valid bugs could be overwhelmed, not helped.
- Severity inflation. Anthropic admits some ratings are inflated. Over time, inflated ratings could erode trust in the labels.
- Disclosure safety. Fast track reports skip human review. How Anthropic stores, transmits and protects unreviewed exploit details, before patches ship, matters a lot.
- Independent numbers. The 88 percent validation rate and the testimonials come from Anthropic. An outside audit of real world accuracy would be valuable.
- Capacity. Anthropic has not said how many projects it can scan, or how often.
Our take
OSS Scanner is a meaningful step. Not because AI bug hunting is new, but because a major lab is now confident enough to remove the human from the loop and say so in public. The early evidence, 72 valid reports out of 74 for wolfSSL and one false positive in 97 checked findings, suggests the confidence is not baseless.
For maintainers of critical projects, the calculation looks favorable: free, repeated scans with reproducers and patches, and the option to stay on the slower human verified path if they prefer. For everyone else, the takeaway is broader. Finding bugs is becoming cheap. Fixing them quickly is now the scarce resource, and that is where the next phase of software security will be decided.
Sources: Anthropic announcement, The Verge.
Source: anthropic.com