Anthropic is trying to solve the difficult balance between using AI for cybersecurity and preventing misuse. The same model that helps security teams fix vulnerabilities can also help attackers break into systems. Its answer, announced this week, is a revamped Cyber Verification Program with three access tiers instead of a simple on-off switch.
The company puts the problem plainly: standard models like Claude Opus 5.5, Claude Fable 5.1, and Claude Sonnet 5.5 ship with conservative cyber safeguards that block most offensive work by default. That’s the safe setting for the general public. It’s also exactly the setting that makes life harder for legitimate security researchers trying to do their jobs.
“The program now consists of three access tiers, which allow security teams to apply for the level of access that best suits their work.” reads the announcement. “Each tier includes access to our most capable models, including Claude Opus 5.5, Claude Sonnet 5.5, Claude Mythos 5.1, and new models moving forward.”
The new program merges two things Anthropic had been running separately, Project Glasswing and the original CVP, into one structure. Defense Access sits at the bottom, covering security operations, incident response, malware reverse engineering, and vulnerability validation. The bar here is low enough that regional hospitals, municipal utilities, open-source maintainers, and individual researchers with a track record can all qualify, and Anthropic says it aims to respond to these applications within days.
Red Team Access steps up to authorized penetration testing, open only to organizations testing systems they’re actually cleared to test. Even here, real-time blocks stay in place for anything that could cause physical harm or mass disruption, deploying ransomware or hitting high-risk safety systems included. Applications at this tier take weeks to review, not days, which tracks given what’s being unlocked.
Specialized Access is reserved for organizations cleared to test systems where a mistake could genuinely hurt people or break markets, things like flight systems, power grids, and interbank transfer infrastructure. Every applicant here gets reviewed in depth alongside the US government, and existing Project Glasswing members simply roll over into this tier without reapplying. That’s about as tightly controlled as commercial AI access gets right now.
Anthropic tested Claude Opus 5.5 with its internal CyScenarioBench benchmark, which measures how well the model can plan and carry out realistic multi-step cyber operations. Without CVP access, all tasks were blocked at the first prompt. In the Red Team Access tier, however, Claude Opus 5.5 completed 34 of 50 tasks without any blocks.
That’s the same completion rate the model hits with no safeguards at all, which is exactly what Specialized Access is meant to represent. The Defense Access tier landed in between, blocking 46 of 50 trials while letting four genuinely difficult ones through. In plain terms, the tiers behave the way they’re supposed to, tight for the general public, open for the people who’ve proven they need the access.

This isn’t just a policy reshuffle, there’s a track record behind it. Through Project Glasswing, Anthropic says partners uncovered at least 129,000 verified software vulnerabilities between April and July 2026, with more than 33,000 rated critical or high severity.
“Through the program, our partners uncovered at least 129,000 verified software vulnerabilities between April and July 2026.” states the announcement. “And through our own open-source scanning efforts, we found an additional 5,500 verified software vulnerabilities between April and October 2026. Of these verified vulnerabilities, more than 33,000 have so far been rated as critical- or high-severity. “

The company is upfront that these numbers are almost certainly a floor, not a ceiling, since the data comes from partial survey responses across a subset of partners. Several partners reportedly told Anthropic the models cut what would have taken months or years down to a fraction of that. If even a conservative read of that holds up, it’s a meaningful shift in how fast defenders can move relative to attackers.
What stands out is not simply that Anthropic built a more capable model, but that it created a verification system designed to match access with the level of risk. Dual-use AI tools need more than a single gate. They need controls that become stricter as the potential impact of misuse increases. The real test will be how this framework performs as more security firms start using it, not just the benchmark results Anthropic published.
Follow me on Twitter: @securityaffairs and Facebook and Mastodon
(SecurityAffairs – hacking, Anthropic)