The Guardrail Paradox: From Full Disclosure To Full Access
Chris Wysopal is Founder and Chief Security Evangelist at Veracode.
gettyIn July 2026, it finally happened.
During an OpenAI red team exercise, one of their AI agents broke out of its test environment. In the first-known frontier model breakout, the model compromised parts of Hugging Face’s infrastructure.
But while the attack is what made headlines across the world, the strangest part of the story took place during the clean-up. When Hugging Face’s responders fed real exploit payloads and C2 (command and control) artifacts into commercial frontier models to reconstruct the attack, they hit a brick wall. Safety guardrails blocked every query because they could not tell the difference between an attacker and an incident responder. As a result, the Hugging Face team had to lean on a self-hosted Chinese model. So, while the attacking agent—the real threat—did not have to answer to a usage policy or ask anyone’s permission, the defenders did.
This is a problem, and we’ve seen it before.
Back in the 1990s, the industry was at odds over vulnerability disclosures. Security researchers looking for flaws in commercial software wanted to publish findings as soon as possible, while vendors argued that publishing technical details was essentially handing attackers a map.
I lived through that argument firsthand as a part of L0pht, a group publicly releasing vulnerabilities we found in Microsoft products. We were young, we worked hard to surface these threats and we were eager to share them with the security community.
Microsoft’s Security Response Center was brand new at the time, and its first director, Scott Culp, reached out and asked whether we would send our findings to Microsoft first and hold our release until after a patch existed. The idea was to balance the public’s right to know about a vulnerability against the need for people to actually patch it before attackers figured out how to exploit it. Despite our “move fast and fix things” ideology, we were able to understand the need to meet Microsoft halfway.
A similar compromise took hold in the broader security industry, and it eventually settled into a workable framework: coordinated vulnerability disclosure, CVEs (common vulnerabilities and exposures), bug bounty programs, and a broader understanding that defenders needed access to the same technical information attackers already had. The approach now sits at the foundation of how the industry manages risk: give the public access to important security information while giving users a chance to protect themselves once attackers learn about the flaw.
That’s the same balance defenders need now when it comes to AI models. Attackers are likely to use whatever model gives them the best results, and they won’t stop to ask if it’s allowed. We saw that play out directly in the Hugging Face case.
Incident response is all about speed. When an intrusion is active, security teams need tools that respond in minutes, not after an approval workflow clears. Some commercial AI providers have built trusted-access programs for exactly this kind of scenario, and they help. But a program that requires enrollment in advance and manual review doesn’t stand a chance against an attack that starts on a Friday night. There are open-weight models already starting to fill that space, particularly for organizations that don’t have the resources to build or run a private frontier model of their own.
The frontier labs themselves have some work to do here, too. They should spend more time on how they test their models for these failure modes. The same goes for how quickly a lab contains an incident once one starts, and how fast it tells the rest of the industry when something operates outside its intended boundaries. This information is best shared in real time when the threat is imminent, not reported on weeks later.
We figured out how to build reporting and notification into vulnerability disclosure. AI incidents need some of that same structure. It’s not realistic to require every lab to open its doors to every researcher, but there are practical ways to get started. For example, defenders under active attack could be pre-approved for expedited use of frontier models, with a straightforward protocol for notifying a lab that they’re responding to a live incident so guardrails don’t get in the way of the work.
Guardrails are necessary because powerful models can be misused. That’s real risk that cannot be ignored. But the guardrail that can’t tell an attacker probing a system from a responder trying to understand what just happened isn’t protecting anyone. When you add friction at the exact moment where delays are the most costly and impactful, you’re putting yourself in an impossible spot.
I’m optimistic that our industry will move quickly on this, because we’ve already learned the lesson with software vulnerabilities. Keeping defenders at arm’s length from technical detail never stopped attackers from getting it themselves. It just leaves the people protecting networks a step behind. The choices frontier labs and policymakers make in the next year or two will shape how incident responders do their jobs the next time a model escapes its boundaries.
The one thing we know for sure is that there will be a next time. What we need to figure out is whether defenders will have the tools they need when it happens.
Forbes Technology Council is an invitation-only community for world-class CIOs, CTOs and technology executives. Do I qualify?
/file/attachments/2998/GroundUp-Gauteng-beautyfunding_572330.jpeg)
