{"id":64554,"date":"2026-07-24T06:23:34","date_gmt":"2026-07-24T10:23:34","guid":{"rendered":"https:\/\/overcentral.com\/en\/?p=64554"},"modified":"2026-07-24T06:23:34","modified_gmt":"2026-07-24T10:23:34","slug":"ai-guardrails-cybersecurity-researchers","status":"publish","type":"post","link":"https:\/\/overcentral.com\/en\/ai-guardrails-cybersecurity-researchers\/","title":{"rendered":"AI Guardrails Push Cybersecurity Researchers to Foreign Models"},"content":{"rendered":"<p>When AI companies erected guardrails to block malicious hackers, they did not anticipate the unintended consequence: driving legitimate cybersecurity researchers toward foreign, ungoverned <a href=\"https:\/\/overcentral.com\/en\/dangerous-ai-models-regulatory-gaps\/\" title=\"Dangerous AI models bypass regulatory controls\" data-iacss-internal=\"1\">AI models<\/a>. For months, the developers of the most advanced frontier models\u2014companies like <a href=\"https:\/\/www.anthropic.com\" target=\"_blank\" rel=\"noopener noreferrer\" data-iacss-external=\"1\">Anthropic<\/a> and <a href=\"https:\/\/www.openai.com\" target=\"_blank\" rel=\"noopener noreferrer\" data-iacss-external=\"1\">OpenAI<\/a>\u2014have rolled out tightly controlled access programs, strict safety filters, and even government-backed export restrictions, all meant to prevent their technology from being weaponized. Yet a growing chorus of network defenders and offensive security researchers now argues that these same barriers are actively harming the very people the industry depends on to find and fix critical vulnerabilities before criminals do.<\/p>\n<h2>How AI Guardrails Are Catching the Wrong People<\/h2>\n<p>The logic behind AI guardrails is straightforward: prevent bad actors from using large language models to write malware, find exploits, or plan cyberattacks. Anthropic, for example, has repeatedly marketed its Mythos model as a kind of doomsday cybermachine, accessible only to carefully vetted users and subject to strict usage constraints. In June, the U.S. government added export control restrictions on both Anthropic\u2019s Mythos and Fable models after reports surfaced that their guardrails could be bypassed. Although the restrictions have since been partially lifted, the underlying philosophy of gatekeeping remains embedded in the industry\u2019s approach to cybersecurity.<\/p>\n<p>Both Anthropic and OpenAI run vetted access programs for researchers\u2014Anthropic\u2019s Cyber Verification Program and OpenAI\u2019s Trusted Access for Cyber program\u2014that offer approved users models with fewer cybersecurity restrictions. The intention is to enable good-faith security work while denying capabilities to attackers. But in practice, the design of these guardrails often fails to distinguish between a defender trying to patch a dangerous bug and an attacker seeking to weaponize it.<\/p>\n<h3>The Blurred Line Between Offense and Defense<\/h3>\n<p>Chris Anley, chief scientist at the security consulting firm NCC Group, articulated the core dilemma: a prompt like \u201cfix this code\u201d is simultaneously an essential defensive mechanism and a roadmap for finding critical vulnerabilities. Anley describes the AI tool as a hammer\u2014impossible to build a house without, but also irreducibly a weapon. When a guardrail blocks the model from responding to such a request, it does not stop attackers; it stops the people trying to secure the foundation.<\/p>\n<p>This ambiguity frustrates researchers whose work depends on probing systems for weaknesses. Mark Dowd, a veteran security researcher known for discovering and selling zero-day vulnerabilities to Western governments, has publicly stated that large, arbitrary decisions about what is safe in security should not be made by random companies. \u201cIt\u2019s not really comfortable to me,\u201d Dowd said during a cybersecurity podcast, reflecting the industry\u2019s growing unease with private AI firms acting as unilateral arbiters of acceptable security research.<\/p>\n<h2>The Practical Cost: Time Spent Negotiating, Not Working<\/h2>\n<p>For researchers who do gain access to the vetted programs, the experience is often inconsistent and frustrating. Chris Thompson, CEO of cybersecurity firm RemoteThreat and founder of Offensive AI Con, reports that guardrails on frontier models can vary from day to day, sometimes over-sanitizing outputs or refusing legitimate requests without clear cause. Even inside the looser boundaries of Anthropic\u2019s and OpenAI\u2019s vetting programs, the inconsistency forces researchers to spend more time negotiating with the model than analyzing vulnerabilities.<\/p>\n<p>\u201cInstead of analyzing a vulnerability and reasoning through the exploitability, you\u2019re trying to find why you\u2019re getting inconsistent results or why are models over-sanitizing the output,\u201d Thompson explained. The practical impact is measurable: precious hours are wasted on coaxing the tool to cooperate rather than on the actual work of securing networks and systems.<\/p>\n<h2>What Is the Cyber Verification Program, and Why Does It Fall Short?<\/h2>\n<p>Anthropic\u2019s Cyber Verification Program is a structured pathway for cybersecurity professionals to gain access to models with relaxed guardrails. Applicants undergo a review process, and if approved, can use Claude Opus and Sonnet for tasks that would normally be blocked, such as analyzing exploit code or assessing vulnerability impact. OpenAI offers a parallel structure through its Trusted Access for Cyber program.<\/p>\n<p>Yet these programs remain narrow in scope and access. One researcher at a smartphone-component manufacturer, who spoke on condition of anonymity, reported that his employer is not part of Anthropic\u2019s program and that the tools are nearly useless for vulnerability discovery as a result. \u201cIf it catches wind we\u2019re doing anything security related, it just stops and isn\u2019t usable,\u201d he said. The gatekeeping effectively locks out legitimate security teams in the very sectors\u2014hardware, manufacturing, critical infrastructure\u2014that most need AI-powered defenses.<\/p>\n<h2>Pushed Toward Foreign Open Source Models<\/h2>\n<p>The most significant consequence of strict guardrails may be the migration of skilled researchers to AI systems that operate entirely outside Western regulatory oversight. Thompson notes that researchers are increasingly turning to Chinese open source models like <a href=\"https:\/\/overcentral.com\/en\/glm-5-2-matches-mythos-cybersecurity\/\" title=\"Z.ai GLM-5.2 Matches Mythos on Cybersecurity\" data-iacss-internal=\"1\">GLM<\/a>aaaa\u2014freely downloadable, locally runnable, and entirely free of vetting or usage restrictions. These models impose no guardrails at all, offering full flexibility for both defensive and offensive work.<\/p>\n<p>\u201cYou have these responsible researchers that are being pushed away from U.S.-governed systems to foreign-owned systems,\u201d Thompson said. \u201cI think it\u2019s more harmful than good to have these guardrails in place.\u201d The irony is acute: by attempting to lock down American AI models, the industry may be handing a strategic advantage to foreign AI ecosystems, which now attract some of the world\u2019s most capable cybersecurity talent.<\/p>\n<h3>Why Offensive Researchers Run Models Locally<\/h3>\n<p>Paolo Stagno, CTO of Crowdfense\u2014a company that develops, acquires, and sells unknown vulnerabilities to government agencies\u2014explained that his team uses frontier AI models only for reverse engineering. When it comes to finding vulnerabilities or building exploits, they avoid cloud-based models entirely out of fear that sensitive vulnerability data could leak or be absorbed into future training runs. Instead, they run open source models locally, where no data leaves the machine and no external entity governs their use.<\/p>\n<p>This shift is not merely a preference; it is a security-driven necessity. Stagno said AI companies treat customers \u201clike children who need babysitting,\u201d and the result is that high-stakes vulnerability research is increasingly conducted on ungoverned foreign models that operate outside any safety framework.<\/p>\n<h2>The Offensive Security Researcher\u2019s Perspective on Guardrails<\/h2>\n<p>Not all security researchers feel impeded by guardrails. Giuseppe Cali, who finds zero-days and develops exploits, said the restrictions do not hinder his workflow because he does not use AI for offensive work in the first place. Instead, he relies on AI tools for initial reverse engineering and code understanding\u2014tasks that fall comfortably within the bounds of most guardrails. \u201cI still want to own the actual bug discovery and weaponization myself,\u201d Cali said. \u201cI am jealous of my bugs, and I like this game too much to let models play it for me.\u201d<\/p>\n<p>Yet Cali\u2019s confidence underscores a broader point: researchers are actively choosing to avoid AI for the sensitive parts of their work, not because of technical limitations, but because the safety mechanisms that govern frontier models create unacceptable risks around data privacy and process transparency. The models are being used less, not more, for the very cybersecurity tasks they were built to accelerate.<\/p>\n<h2>An Industry at a Crossroads<\/h2>\n<p>The tension between AI safety and cybersecurity effectiveness is not theoretical. When the U.S. government slapped export controls on Anthropic\u2019s models in June, the move was prompted by a report that claimed it was possible to bypass their guardrails for malicious purposes. Whether the ban was truly motivated by jailbreak fears remains contested, but the outcome was clear: legitimate researchers lost access to powerful tools, and no equivalent security gap was closed.<\/p>\n<p>Thompson warned that the current trajectory is unsustainable. Rather than tightening restrictions further, he called for AI frontier labs to open up their programs, provide responsible access, and hold those who abuse their tools accountable. \u201cThere\u2019s this big storm coming. There\u2019s this big wave of attacks that are going to happen at speed and scale like never before,\u201d he said. \u201cBut the same security consulting firms and legit researchers that are trying to make a difference are being stifled right now.\u201d<\/p>\n<p>The paradox is that the very companies best positioned to help defend against <a href=\"https:\/\/overcentral.com\/en\/chinese-llms-attacker-defender-gap\/\" title=\"Chinese LLMs Broaden the Gap Between Attackers and Defenders\" data-iacss-internal=\"1\">AI-powered attacks<\/a> are being pushed away from the best AI tools. When guardrails block defenders from doing their jobs, the tools that remain available are the ones with no guardrails at all\u2014running on foreign models, outside any safety or accountability framework. The cybersecurity industry must reckon with a difficult question: are guardrails protecting the internet, or are they simply handing the advantage to those who operate beyond their reach?<\/p>\n","protected":false},"excerpt":{"rendered":"<p>When AI companies erected guardrails to block malicious hackers, they did not anticipate the unintended consequence: driving legitimate cybersecurity researchers toward foreign, ungoverned AI models. For months, the developers of the most advanced frontier models\u2014companies like Anthropic and OpenAI\u2014have rolled out tightly controlled access programs, strict safety filters, and even government-backed export restrictions, all meant [&hellip;]<\/p>\n","protected":false},"author":7,"featured_media":83768,"comment_status":"closed","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"fifu_image_url":"https:\/\/cards.overcentral.com\/cards\/en\/64554.png","fifu_image_alt":"AI Guardrails Push Cybersecurity Researchers to Foreign Models","footnotes":""},"categories":[349],"tags":[],"class_list":["post-64554","post","type-post","status-publish","format-standard","has-post-thumbnail","category-articles"],"fifu_image_url":"https:\/\/cards.overcentral.com\/cards\/en\/64554.png","fifu_image_alt":"AI Guardrails Push Cybersecurity Researchers to Foreign Models","_links":{"self":[{"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/posts\/64554","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/users\/7"}],"replies":[{"embeddable":true,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/comments?post=64554"}],"version-history":[{"count":0,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/posts\/64554\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/media\/83768"}],"wp:attachment":[{"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/media?parent=64554"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/categories?post=64554"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/tags?post=64554"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}