AI Guardrails Push Cybersecurity Researchers to Foreign Models

AI safety guardrails designed to block attackers are inadvertently driving legitimate cybersecurity researchers to use ungoverned foreign AI models.

By Central
The article explores how AI guardrails intended for safety are pushing cybersecurity researchers to seek ungoverned foreign models.
Highlights
  • AI guardrails block legitimate cybersecurity researchers from probing vulnerabilities, forcing them to use foreign models.
  • Chris Anley of NCC Group describes the AI tool as a hammer that is both essential for building and a weapon.
  • The U.S. government's export controls on Anthropic's models further restricted access for legitimate researchers.

When AI companies erected guardrails to block malicious hackers, they did not anticipate the unintended consequence: driving legitimate cybersecurity researchers toward foreign, ungoverned AI models. For months, the developers of the most advanced frontier models—companies like Anthropic and OpenAI—have rolled out tightly controlled access programs, strict safety filters, and even government-backed export restrictions, all meant to prevent their technology from being weaponized. Yet a growing chorus of network defenders and offensive security researchers now argues that these same barriers are actively harming the very people the industry depends on to find and fix critical vulnerabilities before criminals do.

How AI Guardrails Are Catching the Wrong People

The logic behind AI guardrails is straightforward: prevent bad actors from using large language models to write malware, find exploits, or plan cyberattacks. Anthropic, for example, has repeatedly marketed its Mythos model as a kind of doomsday cybermachine, accessible only to carefully vetted users and subject to strict usage constraints. In June, the U.S. government added export control restrictions on both Anthropic’s Mythos and Fable models after reports surfaced that their guardrails could be bypassed. Although the restrictions have since been partially lifted, the underlying philosophy of gatekeeping remains embedded in the industry’s approach to cybersecurity.

Both Anthropic and OpenAI run vetted access programs for researchers—Anthropic’s Cyber Verification Program and OpenAI’s Trusted Access for Cyber program—that offer approved users models with fewer cybersecurity restrictions. The intention is to enable good-faith security work while denying capabilities to attackers. But in practice, the design of these guardrails often fails to distinguish between a defender trying to patch a dangerous bug and an attacker seeking to weaponize it.

The Blurred Line Between Offense and Defense

Chris Anley, chief scientist at the security consulting firm NCC Group, articulated the core dilemma: a prompt like “fix this code” is simultaneously an essential defensive mechanism and a roadmap for finding critical vulnerabilities. Anley describes the AI tool as a hammer—impossible to build a house without, but also irreducibly a weapon. When a guardrail blocks the model from responding to such a request, it does not stop attackers; it stops the people trying to secure the foundation.

This ambiguity frustrates researchers whose work depends on probing systems for weaknesses. Mark Dowd, a veteran security researcher known for discovering and selling zero-day vulnerabilities to Western governments, has publicly stated that large, arbitrary decisions about what is safe in security should not be made by random companies. “It’s not really comfortable to me,” Dowd said during a cybersecurity podcast, reflecting the industry’s growing unease with private AI firms acting as unilateral arbiters of acceptable security research.

The Practical Cost: Time Spent Negotiating, Not Working

For researchers who do gain access to the vetted programs, the experience is often inconsistent and frustrating. Chris Thompson, CEO of cybersecurity firm RemoteThreat and founder of Offensive AI Con, reports that guardrails on frontier models can vary from day to day, sometimes over-sanitizing outputs or refusing legitimate requests without clear cause. Even inside the looser boundaries of Anthropic’s and OpenAI’s vetting programs, the inconsistency forces researchers to spend more time negotiating with the model than analyzing vulnerabilities.

“Instead of analyzing a vulnerability and reasoning through the exploitability, you’re trying to find why you’re getting inconsistent results or why are models over-sanitizing the output,” Thompson explained. The practical impact is measurable: precious hours are wasted on coaxing the tool to cooperate rather than on the actual work of securing networks and systems.

What Is the Cyber Verification Program, and Why Does It Fall Short?

Anthropic’s Cyber Verification Program is a structured pathway for cybersecurity professionals to gain access to models with relaxed guardrails. Applicants undergo a review process, and if approved, can use Claude Opus and Sonnet for tasks that would normally be blocked, such as analyzing exploit code or assessing vulnerability impact. OpenAI offers a parallel structure through its Trusted Access for Cyber program.

Yet these programs remain narrow in scope and access. One researcher at a smartphone-component manufacturer, who spoke on condition of anonymity, reported that his employer is not part of Anthropic’s program and that the tools are nearly useless for vulnerability discovery as a result. “If it catches wind we’re doing anything security related, it just stops and isn’t usable,” he said. The gatekeeping effectively locks out legitimate security teams in the very sectors—hardware, manufacturing, critical infrastructure—that most need AI-powered defenses.

Pushed Toward Foreign Open Source Models

The most significant consequence of strict guardrails may be the migration of skilled researchers to AI systems that operate entirely outside Western regulatory oversight. Thompson notes that researchers are increasingly turning to Chinese open source models like GLMaaaa—freely downloadable, locally runnable, and entirely free of vetting or usage restrictions. These models impose no guardrails at all, offering full flexibility for both defensive and offensive work.

“You have these responsible researchers that are being pushed away from U.S.-governed systems to foreign-owned systems,” Thompson said. “I think it’s more harmful than good to have these guardrails in place.” The irony is acute: by attempting to lock down American AI models, the industry may be handing a strategic advantage to foreign AI ecosystems, which now attract some of the world’s most capable cybersecurity talent.

Why Offensive Researchers Run Models Locally

Paolo Stagno, CTO of Crowdfense—a company that develops, acquires, and sells unknown vulnerabilities to government agencies—explained that his team uses frontier AI models only for reverse engineering. When it comes to finding vulnerabilities or building exploits, they avoid cloud-based models entirely out of fear that sensitive vulnerability data could leak or be absorbed into future training runs. Instead, they run open source models locally, where no data leaves the machine and no external entity governs their use.

This shift is not merely a preference; it is a security-driven necessity. Stagno said AI companies treat customers “like children who need babysitting,” and the result is that high-stakes vulnerability research is increasingly conducted on ungoverned foreign models that operate outside any safety framework.

The Offensive Security Researcher’s Perspective on Guardrails

Not all security researchers feel impeded by guardrails. Giuseppe Cali, who finds zero-days and develops exploits, said the restrictions do not hinder his workflow because he does not use AI for offensive work in the first place. Instead, he relies on AI tools for initial reverse engineering and code understanding—tasks that fall comfortably within the bounds of most guardrails. “I still want to own the actual bug discovery and weaponization myself,” Cali said. “I am jealous of my bugs, and I like this game too much to let models play it for me.”

Yet Cali’s confidence underscores a broader point: researchers are actively choosing to avoid AI for the sensitive parts of their work, not because of technical limitations, but because the safety mechanisms that govern frontier models create unacceptable risks around data privacy and process transparency. The models are being used less, not more, for the very cybersecurity tasks they were built to accelerate.

An Industry at a Crossroads

The tension between AI safety and cybersecurity effectiveness is not theoretical. When the U.S. government slapped export controls on Anthropic’s models in June, the move was prompted by a report that claimed it was possible to bypass their guardrails for malicious purposes. Whether the ban was truly motivated by jailbreak fears remains contested, but the outcome was clear: legitimate researchers lost access to powerful tools, and no equivalent security gap was closed.

Thompson warned that the current trajectory is unsustainable. Rather than tightening restrictions further, he called for AI frontier labs to open up their programs, provide responsible access, and hold those who abuse their tools accountable. “There’s this big storm coming. There’s this big wave of attacks that are going to happen at speed and scale like never before,” he said. “But the same security consulting firms and legit researchers that are trying to make a difference are being stifled right now.”

The paradox is that the very companies best positioned to help defend against AI-powered attacks are being pushed away from the best AI tools. When guardrails block defenders from doing their jobs, the tools that remain available are the ones with no guardrails at all—running on foreign models, outside any safety or accountability framework. The cybersecurity industry must reckon with a difficult question: are guardrails protecting the internet, or are they simply handing the advantage to those who operate beyond their reach?

Share This Article