Industry

AI guardrails hinder offensive cybersecurity researchers' work

AI companies like Anthropic and OpenAI have implemented strict guardrails and vetted programs to prevent malicious use of their models. However, these restrictions are now impeding legitimate offensive cybersecurity researchers who need to find vulnerabilities and develop exploits. Researchers report spending more time negotiating with models than working on security, and some are turning to open-source or Chinese models without restrictions.

Neura News

Neura News

Neura Market Editorial

July 24, 20266 min read
AI guardrails hinder offensive cybersecurity researchers' work

For months, major AI companies have created special vetted programs and strict guardrails to limit how malicious hackers can use their models. But these limits are now getting in the way of legitimate network defenders and offensive cybersecurity researchers.

Regardless of whether the incident was truly motivated by fears of a jailbreak, Anthropic has repeatedly marketed Mythos as a kind of doomsday cybermachine that can only be given to carefully vetted users, and even then with strict guardrails in place. (The export controls on Fable 5 and Mythos 5 have since been lifted. Fable 5 returned to general access on July 1. Mythos 5 has been reintroduced only to vetted U.S. organizations as part of the government's review process.)

That kind of gatekeeping is not unique to Mythos. Both Anthropic, with its other models, and OpenAI offer cybersecurity researchers programs they can apply to in order to get vetted and, if approved, access models with fewer cybersecurity restrictions. These programs are OpenAI's Trusted Access for Cyber and Anthropic's Cyber Verification Program.

These guardrails have been widely criticized, especially by researchers whose job is to find unknown vulnerabilities in systems and devise ways to exploit them before criminals do.

Researchers speak out

During a recent appearance on a cybersecurity podcast, Mark Dowd, a well-known security researcher, said that "it's not really comfortable to me that these random large companies are making arbitrary decisions about what is safe in security and what's not."

Dowd has spent decades finding and selling "zero days" (previously unknown software flaws and the exploits that take advantage of them) to Western governments, rather than reporting them to the software makers so they get patched. Governments pay a premium for vulnerabilities precisely because they stay open, which is useful for intelligence operations.

Dowd admitted his work may make him biased, but he is not alone. Several people who work in offensive cybersecurity (they proactively probe systems for weaknesses) described to TechCrunch how they use AI tools and deal with their guardrails.

Chris Anley, the chief scientist at security consulting giant NCC Group, said that asking an AI model to try to exploit a bug is a key step in confirming it is a real vulnerability worth fixing. But if a guardrail prompts the model to refuse to answer the question outright, the guardrail hurts defenders, he said.

"This is where the whole offensive versus defensive and guardrails part comes in, because 'fix this code' as a prompt is both an essential mechanism for defense but also a roadmap for finding critical vulnerabilities in the code base," said Anley. "So at the same time, the same tool is both an offensive tool and a defensive tool, and the two can't really be unpicked."

It is "like a hammer," he continued. "You can't build a house without a hammer. It's definitely a tool but it's also irreducibly a weapon as well."

When he and his colleagues run into such a roadblock, they sometimes fall back on open-source AI models that come with no guardrails at all.

Paolo Stagno, the chief technology officer at CrowdFense, a well-known company that develops, acquires, and sells unknown vulnerabilities to government agencies, agreed with Dowd. He said AI companies "essentially treat customers like children who need babysitting" with their vetted programs and guardrails.

Stagno said he and his colleagues do use frontier models, but only for reverse engineering. They avoid using AI to help find vulnerabilities or build exploits, he said, because feeding that work into a cloud-based model risks leaking sensitive vulnerability data or having it absorbed into future training runs. For that step, he said, they use open source models run locally, as they do not rely on sharing data outside of the model.

The #1 Newsletter in AI

Stay ahead of the AI curve

The most important updates, news, and content — delivered weekly.

No spam. Unsubscribe anytime.

Different approaches to AI use

Giuseppe Cali, a security researcher who finds zero-days and develops exploits, said guardrails are not impeding his work. That is because he does not use AI for offensive work. Instead, he uses it for initial reverse engineering, to understand the code he is analyzing, and to build supporting tools. For that, he said, AI tools can speed up the process and allow him to focus on discovering vulnerabilities.

"I still want to own the actual bug discovery and weaponization myself and that wouldn't change if all guardrails were lifted tomorrow," said Cali. "I am jealous of my bugs, and I like this game too much to let models play it for me."

One researcher at a smartphone-component manufacturer, who spoke on condition of anonymity because he is not authorized to talk to the press, said his employer is not part of Anthropic's CVP program. As a result, its tools are barely useful for finding vulnerabilities because the guardrails are too strict.

"If it catches wind we're doing anything security related, it just stops and isn't usable," the person said.

Chris Thompson, chief executive of cybersecurity firm RemoteThreat and founder of Offensive AI Con, an offensive security and AI-focused event, said that in his experience using the frontier AI models, the guardrails can be inconsistent and work differently every day. That is true even inside the looser boundaries of Anthropic and OpenAI's vetted programs.

"I think the practical impact is you spend a lot of time negotiating with the model instead of working on the core security program," said Thompson. "Instead of analyzing a vulnerability and reasoning through the exploitability, you're trying to find why you're getting inconsistent results or why are models over-sanitizing the output."

Shift to foreign models

As a consequence, researchers rely on or get pushed toward Chinese open-source models like GLM (freely downloadable models that can be run locally with no vetting or usage restrictions), said Thompson.

"You have these responsible researchers that are being pushed away from U.S.-governed systems to foreign-owned systems," he said. "I think it's more harmful than good to have these guardrails in place."

Rather than tightening restrictions further, Thompson called for the AI frontier labs to open up their programs, provide responsible access, and also hold those who abuse their tools accountable. Otherwise, he argued, defenders will lose the AI race.

"There's this big storm coming. There's this big wave of attacks that are going to happen at speed and scale like never before," said Thompson. "But the same security consulting firms and legit researchers that are trying to make a difference are being stifled right now."

Related on Neura Market

More from Neura News

Industry

Power line failure reveals AI data center grid risks and solutions

A fallen power line near Washington, DC caused over 3 gigawatts of data center load to vanish from the PJM grid in seconds, spiking voltage across the region. The event, which made lights flicker from Northern Virginia to Chicago, highlights a growing problem as AI data centers become larger and more concentrated. Experts warn that without better coordination or technology like ON.Energy's battery-backed uninterruptible power supply, such disruptions will become more frequent and severe.

Jul 25·5 min read