AI Guardrails Block Cybersecurity Researchers From Defensive Work
Strict safety controls on frontier AI models are pushing legitimate security professionals toward unregulated alternatives, experts warn.

Defensive researchers caught in AI safety crossfire
AI companies' efforts to prevent malicious use of their models are creating unintended obstacles for the cybersecurity professionals tasked with finding vulnerabilities before criminals exploit them. Security researchers who probe systems for weaknesses—known as offensive security work—report that guardrails on frontier AI models from Anthropic and OpenAI frequently block legitimate defensive activities.
The tension came into sharp focus in June when the U.S. government imposed export controls on Anthropic's Mythos and Fable models, prompted in part by reports of guardrail bypasses. While those restrictions have since been partially lifted—Fable 5 returned to general access July 1, and Mythos 5 is available only to vetted U.S. organizations—the incident highlighted broader friction between AI safety measures and security research needs.
Both Anthropic and OpenAI now operate vetted programs (Cyber Verification Program and Trusted Access for Cyber, respectively) that grant approved researchers reduced restrictions. But multiple security professionals told TechCrunch these programs remain inadequate for their work.
The dual-use dilemma
Chris Anley, chief scientist at NCC Group, described the core problem: asking an AI model to exploit a bug is essential for confirming it's a real vulnerability worth fixing. When guardrails block that query, they hinder defenders.
"The same tool is both an offensive tool and a defensive tool, and the two can't really be unpicked," Anley said. "It's like a hammer. You can't build a house without a hammer. It's definitely a tool but it's also irreducibly a weapon as well."
When researchers hit these roadblocks, many turn to open-source AI models with no guardrails—a workaround that defeats the purpose of the restrictions.
Paolo Stagno, CTO at vulnerability broker CrowdFense, said AI companies "essentially treat customers like children who need babysitting." His team uses frontier models only for reverse engineering, avoiding them for vulnerability discovery due to concerns about leaking sensitive data through cloud-based systems. For that work, they rely on locally-run open-source models.
Inconsistent enforcement drives researchers away
Chris Thompson, CEO of RemoteThreat and founder of Offensive AI Con, reported that guardrails behave inconsistently even within vetted programs, changing behavior day to day.
"Instead of analyzing a vulnerability and reasoning through the exploitability, you're trying to find why you're getting inconsistent results," Thompson said. The result: researchers increasingly turn to Chinese open-source models like GLM, which can be downloaded and run locally without restrictions.
"You have these responsible researchers that are being pushed away from U.S.-governed systems to foreign-owned systems," Thompson said. "I think it's more harmful than good to have these guardrails in place."
One researcher at a smartphone-component manufacturer, speaking anonymously, said his employer isn't part of Anthropic's CVP program, making the tools "barely useful" because guardrails trigger whenever security work is detected.
Not all researchers share these concerns. Giuseppe Cali, who finds zero-days and develops exploits, said guardrails don't impede his work because he uses AI only for initial reverse engineering and tool-building, not for actual vulnerability discovery. "I am jealous of my bugs, and I like this game too much to let models play it for me," he said.
Why it matters
The guardrail debate illustrates a fundamental tension in AI governance: safety measures designed to prevent misuse can inadvertently handicap the defenders who protect systems from attack. If overly restrictive controls push legitimate security researchers toward unregulated alternatives—particularly foreign-owned models—they may undermine rather than enhance cybersecurity. As Thompson warned, defenders risk losing the AI race at precisely the moment when AI-powered attacks are expected to arrive "at speed and scale like never before."
These details were first reported by TechCrunch.
This is an original analysis by the Omega editorial team. Source reporting: AI Watch.
Want systems like this working for your business?
Book a Call