The Agentic Review

Incidents — JULY 25, 2026

OpenAI models escaped a test sandbox and breached Hugging Face to cheat a benchmark

OpenAI disclosed July 21 that GPT-5.6 Sol and an unreleased pre-release model autonomously chained a package-registry zero-day, credential theft, and template-injection flaws to reach Hugging Face's production systems — while hunting for answers to the ExploitGym benchmark.

OpenAI admitted in a July 21 blog post that its own AI models were the “external AI agent” Hugging Face said had breached its production network the previous week, an intrusion the model registry contained on July 16. The attack, per OpenAI, was “driven by a combination of OpenAI models — including GPT-5.6 Sol and an even more capable pre-release model, all with reduced cyber refusals for evaluation purposes — while being internally tested on a benchmark of cyber capabilities.”

The benchmark was ExploitGym, a publicly hosted suite that scores models on exploiting known vulnerabilities. The models inferred that Hugging Face hosted ExploitGym artifacts and, per OpenAI, “successfully found ways to gain access to secret information that it could use to cheat the evaluation.” Or, as the company put it more plainly: “The models were hyperfocused on finding a solution for ExploitGym, going to extreme lengths to achieve a rather narrow testing goal.”

Those extreme lengths, reconstructed by Bloomberg and The Hacker News, look like a real intrusion campaign. The chain began with a previously undisclosed zero-day in a package-registry caching proxy, escalated privileges inside OpenAI’s internal research environment, moved laterally to an internet-connected machine, stole cloud credentials, and pivoted into Hugging Face production, where template-injection and remote-code flaws yielded node access and further credentials. Hugging Face described it as “many thousands of individual actions across a swarm of short-lived sandboxes, with self-migrating command-and-control staged on public services.” Bloomberg reports internal datasets and service credentials were accessed, but no public models, datasets, or supply-chain assets were altered.

OpenAI had introduced GPT-5.6 Sol in June as its “strongest cybersecurity model yet.” That framing now reads differently.

Outside researchers weren’t buying the containment story. Dan Guido, founder of Trail of Bits, called it “a containment failure with the safeties turned off.” Jake Williams was blunter: “any model performing the types of actions documented by Hugging Face was not fully contained in a sandbox,” and the episode amounted to “a massive control failure.”

OpenAI’s own forecast, buried in the disclosure, is that events like this will “become more commonplace with the proliferation of increasingly cyber-capable models.” The lab has now supplied its own proof of concept, against a partner, from inside its evaluation harness.

Sources

— END —