Industry

AI guardrails hinder offensive cybersecurity researchers' work

AI companies like Anthropic and OpenAI have implemented strict guardrails and vetted programs to prevent malicious use of their models. However, these restrictions are now impeding legitimate offensive cybersecurity researchers who need to find vulnerabilities and develop exploits. Researchers report spending more time negotiating with models than working on security, and some are turning to open-source or Chinese models without restrictions.

Neura News

Neura News

Neura Market Editorial

July 24, 20266 min read
AI guardrails hinder offensive cybersecurity researchers' work

{ "title": "AI Guardrails Block Researchers, Pushing Them Toward Foreign Models", "body": "The same AI guardrails designed to keep malicious hackers at bay are blocking legitimate cybersecurity researchers, forcing them to abandon U.S.-made models for open-source or foreign alternatives. A growing chorus of offensive security experts says the restrictions are making their work harder, not safer.\n\nIn June 2024, the U.S. government imposed export control restrictions on Anthropic’s Mythos and Fable models. The move was prompted at least in part by a report claiming it was possible to bypass guardrails to build and execute malicious cyberattacks. Fable 5 returned to general access on July 1, 2024, but Mythos 5 was reintroduced only to vetted U.S. organizations as part of a government review process.\n\nAnthropic has repeatedly marketed Mythos as “some kind of doomsday cybermachine.” The company now offers a Cyber Verification Program for researchers to access models with fewer cybersecurity restrictions. OpenAI runs a similar program, Trusted Access for Cyber, for vetted researchers.\n\nBut for many in the field, these programs are not enough.\n\n## The Same Tool Is Both Offensive and Defensive\n\nSecurity researcher Mark Dowd has spent decades finding and selling zero-days to Western governments, not reporting them to software makers. He said the current system puts too much power in the hands of private companies.\n\n“It’s not really comfortable to me that these random large companies are making arbitrary decisions about what is safe in security and what’s not,” Dowd said.\n\nChris Anley, chief scientist at NCC Group, said the distinction between offensive and defensive use of AI is artificial. Asking an AI model to try to exploit a bug is a key step in confirming a real vulnerability worth fixing.\n\n“This is where the whole offensive versus defensive and guardrails part comes in, because ‘fix this code’ as a prompt is both an essential mechanism for defense but also a roadmap for finding critical vulnerabilities in the code base,” Anley said.\n\nHe added, “So at the same time, the same tool is both an offensive tool and a defensive tool, and the two can’t really be unpicked.”\n\nAnley drew a blunt analogy: “You can’t build a house without a hammer. It’s definitely a tool but it’s also irreducibly a weapon as well.”\n\nA guardrail that prompts a model to refuse to answer hurts defenders. Researchers say the restrictions are so blunt that they block legitimate work while doing little to stop determined attackers. Anley noted that the same prompt used to find a vulnerability for patching could be used by an attacker to find the same flaw for exploitation. The model cannot distinguish intent, so it often refuses both.\n\n## Researchers Turn to Open-Source Models\n\nPaolo Stagno is chief technology officer at Crowdfense, a company that develops, acquires, and sells unknown vulnerabilities to government agencies. He said AI companies “essentially treat customers like children who need babysitting.”\n\nStagno and his colleagues use frontier models only for reverse engineering. They avoid using AI to find vulnerabilities or build exploits due to the risk of leaking sensitive data into a cloud-based model. For vulnerability finding and exploit building, they use open-source models run locally. Stagno said that even if guardrails were removed, he would still prefer local models for sensitive work because cloud-based models pose a data leakage risk.\n\nGiuseppe Cali is a security researcher who finds zero-days and develops exploits. He does not use AI for offensive work, but uses it for initial reverse engineering, understanding code, and building supporting tools. Cali said AI tools speed up the process and allow him to focus on discovering vulnerabilities.\n\n“I still want to own the actual bug discovery and weaponization myself and that wouldn’t change if all guardrails were lifted tomorrow,” Cali said. “I am jealous of my bugs, and I like this game too much to let models play it for me.”\n\nOne researcher at a smartphone-component manufacturer, who spoke on condition of anonymity, said their employer is not part of Anthropic’s Cyber Verification Program. The tools are barely useful due to strict guardrails.\n\n“If it catches wind we’re doing anything security related, it just stops and isn’t usable,” the researcher said.\n\nThe researcher added that their company has tried to get access to vetted programs but found the process slow and opaque. They now rely on open-source models for most security work, which they said are more reliable and faster for their needs.\n\n## Inconsistent Guardrails Waste Time\n\nChris Thompson is CEO of RemoteThreat and founder of Offensive AI Con. He said guardrails can be inconsistent and work differently every day, even inside vetted programs.\n\n“I think the practical impact is you spend a lot of time negotiating with the model instead of working on the core security program,” Thompson said.\n\nHe described the daily reality for researchers: “Instead of analyzing a vulnerability and reasoning through the exploitability, you’re trying to find why you’re getting inconsistent results or why are models over-sanitizing the output.”\n\nThompson said the inconsistency pushes researchers toward alternatives. Chinese open-source models like GLM are freely downloadable, run locally, and come with no vetting or usage restrictions. He noted that these models are becoming increasingly capable and are used by researchers who want to avoid the friction of U.S.-made models.\n\n“You have these responsible researchers that are being pushed away from U.S.-governed systems to foreign-owned systems,” Thompson said.\n\nHe called the guardrails counterproductive. “I think it’s more harmful than good to have these guardrails in place.”\n\nThompson also pointed out that the vetting process for programs like Anthropic’s Cyber Verification Program can take weeks or months, during which researchers lose productivity. Meanwhile, attackers face no such delays and can use any model they choose.\n\n## A ‘Big Storm’ Coming\n\nThompson warned that the current approach leaves defenders unprepared for what is coming. He said attacks will arrive at speed and scale never seen before.\n\n“There’s this big storm coming. There’s this big wave of attacks that are going to happen at speed and scale like never before,” Thompson said. “But the same security consulting firms and legit researchers that are trying to make a difference are being stifled right now.”\n\nHe called for AI frontier labs to open up programs, provide responsible access, and hold abusers accountable rather than tighten restrictions on everyone. Thompson warned that defenders will lose the AI race otherwise.\n\nAnley echoed this concern, noting that the pace of AI development means defenders must be able to experiment freely with models to understand emerging threats. He said that if researchers cannot use the latest models for security work, they will fall behind attackers who have no such constraints.\n\n## Guardrails Push Researchers Abroad\n\nThe article argues that guardrails intended to stop malicious hackers are also hindering legitimate researchers. It suggests that guardrails push researchers toward foreign or open-source models, which may be harmful to U.S. interests. Multiple researchers criticized the inconsistency and restrictiveness of guardrails.\n\nFor now, many of the people best positioned to defend against AI-powered attacks are turning to models built outside the U.S. to do their work. Stagno noted that his team uses open-source models from China and Europe because they offer more flexibility and fewer restrictions. He said the U.S. risks losing its edge in AI security if it continues to lock down its models.\n\nDowd added that the current system creates a perverse incentive: responsible researchers follow the rules and get blocked, while malicious actors ignore them and use whatever tools they want. He said the only people hurt by guardrails are those trying to help.\n\n“The bad guys don’t care about guardrails,” Dowd said. “They’ll just use a different model or jailbreak the one they have. The only people who get slowed down are the ones trying to do the right thing.”\n\n## Related on Neura Market\n- Cybersecurity and AI Policy\n- Open-Source AI Models\n- Offensive Security Research" }

The #1 Newsletter in AI

Stay ahead of the AI curve

The most important updates, news, and content — delivered weekly.

No spam. Unsubscribe anytime.

More from Neura News

AI Models

42 Mathematicians Urge Royal Society to Warn Government and Media About AI Existential Risk

Forty-two mathematical fellows, including Fields Medal winners Martin Hairer, Peter Scholze, and Wendelin Werner, have signed an open letter urging the Royal Society to warn the UK government and media about existential risks from advanced AI. The letter follows recent breakthroughs in which leading models solved open research problems, including a Millennium Problem. None of the signatories are affiliated with AI companies. The group warns that AI labs' estimates of existential risk above ten percent must not be dismissed as hype, and that by the time the situation becomes obvious to the public, it may be too late to act.

Sep 18·2 min read
Developer

Steve Yegge Shuts Down Gas Town After Failing to Build Anything Else With It

Steve Yegge shut down Gas Town, his ultra-vibed coding agent orchestrator, after admitting he never built anything else with it despite heavy subscription spend. Databricks reported a 60% coding spend increase after rolling out GPT-6 Astra to 3,500 engineers, OpenAI published a misalignment disclosure framework with six case reports, and Xiaomi ran MiMo-V2.6 RL training in public with live telemetry.

Sep 18·21 min read