AI Guardrails Frustrate Cybersecurity Researchers, Experts Say
4 min read
Artificial intelligence has become a valuable tool for cybersecurity professionals, helping them analyze software, identify weaknesses, and speed up security research. However, many offensive cybersecurity researchers now argue that the strict safety controls built into leading AI models are making their jobs harder rather than improving security.
For months, major AI companies have introduced vetted access programs and strict guardrails designed to stop criminals from using advanced AI models to launch cyberattacks. While these measures are meant to improve safety, researchers say they are also slowing down legitimate security work.
AI Safety Measures Come Under Scrutiny
The debate gained attention after the U.S. government imposed export restrictions on Anthropic’s AI models, Mythos and Fable, in June. The decision followed reports claiming that users could bypass the models’ cybersecurity safeguards and potentially use them for malicious activities.
Although it remains unclear whether those jailbreak concerns directly triggered the restrictions, Anthropic has consistently positioned Mythos as a highly capable AI model that should only be available to carefully vetted users with strict limitations.
Since then, the restrictions have been eased. Fable 5 became generally available again on July 1, while Mythos 5 has returned only for approved U.S. organizations participating in the government’s review process.
Security Researchers Face Extra Barriers
Anthropic is not the only company taking this approach. Both Anthropic and OpenAI offer special access programs that allow approved cybersecurity professionals to use AI models with fewer restrictions. These include Anthropic’s Cyber Verification Program and OpenAI’s Trusted Access for Cyber initiative.
Even with these programs in place, many researchers believe the current safeguards are too restrictive.
Veteran security researcher Mark Dowd recently criticized the growing role of AI companies in deciding what cybersecurity activities should or should not be allowed. Speaking on a cybersecurity podcast, Dowd questioned whether large technology companies should have the authority to make what he described as arbitrary decisions about acceptable security research.
Dowd has spent years discovering and selling previously unknown software vulnerabilities, commonly known as zero-days, to Western governments instead of reporting them directly to software vendors for patches. Governments often pay significant amounts for these vulnerabilities because they can be useful for intelligence operations.
Although Dowd acknowledged that his background may influence his opinion, several other offensive cybersecurity experts share similar concerns.
Offensive and Defensive Security Often Overlap
Chris Anley, chief scientist at cybersecurity consulting firm NCC Group, explained that AI plays an important role in confirming whether a software bug is actually exploitable.
According to Anley, if an AI assistant refuses to analyze or test a vulnerability because of built-in guardrails, security teams lose an important tool that helps them verify real security issues before attackers can exploit them.
He noted that prompts asking an AI model to “fix this code” can serve two purposes at once. They help developers strengthen software defenses while also revealing weaknesses that attackers might exploit. Because of this overlap, separating offensive and defensive cybersecurity is much more difficult than AI companies may assume.
Anley compared AI to a hammer, saying it is both an essential construction tool and something that can also be misused as a weapon.
When commercial AI models refuse to assist, his team sometimes switches to open-source AI models that have no built-in restrictions.
Open-Source AI Becomes an Alternative
Paolo Stagno, chief technology officer at Crowdfense, also questioned the strict vetting systems used by AI companies. He argued that many providers treat professional security researchers “like children who need babysitting.”
Stagno said his team uses advanced commercial AI models mainly for reverse engineering software. However, they avoid using cloud-based AI systems for vulnerability discovery or exploit development because they worry sensitive security data could be exposed or incorporated into future model training.
Instead, they rely on open-source AI models running locally, allowing sensitive research to remain entirely within their own systems.
Researchers Use AI Differently
Not every offensive security expert believes guardrails are a major obstacle.
Independent researcher Giuseppe Cali said AI mainly helps him understand complex code and create supporting tools during reverse engineering. He still prefers to discover vulnerabilities and develop exploits himself.
According to Cali, even if every AI restriction disappeared tomorrow, he would continue handling bug discovery personally because he enjoys the research process and wants full ownership of his findings.
Companies Report Inconsistent AI Restrictions
One anonymous researcher working for a smartphone component manufacturer said his employer is not part of Anthropic’s Cyber Verification Program. As a result, he finds the company’s AI tools far less useful for vulnerability research.
He explained that whenever the AI detects security-related work, it often refuses to continue, making it difficult to complete legitimate research tasks.
Chris Thompson, CEO of cybersecurity company RemoteThreat and founder of Offensive AI Con, said another challenge is inconsistency. In his experience, AI guardrails often behave differently from one day to the next—even within approved researcher programs offered by Anthropic and OpenAI.
Rather than focusing entirely on vulnerability analysis, researchers frequently spend valuable time trying to convince AI models to provide consistent responses.
Researchers Turn to Foreign AI Models
Because of these limitations, Thompson said many cybersecurity professionals increasingly rely on open-source Chinese AI models such as GLM. These models can be downloaded, run locally, and used without approval programs or usage restrictions.
He believes this trend could have unintended consequences by encouraging responsible researchers to move away from AI systems developed under U.S. oversight.
Instead of expanding restrictions, Thompson argued that leading AI companies should widen access for verified security researchers while holding users accountable if they misuse the technology.
According to him, cyber threats are expected to grow rapidly as AI becomes more capable, and security professionals need access to the same advanced tools if they hope to keep pace with increasingly sophisticated attacks.
As AI continues transforming cybersecurity, the debate over balancing safety with legitimate research is likely to remain one of the industry’s biggest challenges.
Also read : Stripe and Advent Bid $53.4B to Take Over PayPal
