By Rajwa Quasim
Open AI admitted on Tuesday that some of its own AI systems slipped past their controlled test environment and hacked into Hugging Face, the world’s largest AI model repository, in what the company is called an “unprecedented cyber incident.”
The incident involved Open AI’s GPT-5.6 Soland an “even more capable pre-release model.” Both were being tested on a cybersecurity exercise called ExploitGym, designed to check how well AI can find and take advantage of security weaknesses.
For the test, OpenAI models were operating with “reduced cyber refusals for evaluation purposes” so it could properly measure the model’s ability to conduct cyber-attacks. The company warned that these kinds of incidents may happen more often as the proliferation of cyber-capable models is increasing.
READ: Chinese AI models by Moonshot, Alibaba ramp up pressure on Silicon Valley (July 20, 2026)
OpenAI explained that the models became so fixated on completing tasks that they pushed to extreme length to succeed, regardless of the consequences. It may even manage to break out of its highly isolated sandboxed environment and obtain open internet access by discovering and exploiting a zero-day vulnerability in an unspecified vendor’s software, which acts as a proxy and cache for package registries.
The company explained that “with this access, our models performed a series of privilege escalation and lateral movement actions in our research testing environment until the models reached a node with internet access.”
After getting past the internet access blockade, the models worked out that Hugging Face as the platform storing models, datasets, and solutions for the ExploitGym. It pushed them to search for ways to access confidential data that they could use to cheat on the benchmark. At one stage, the models combined multiple attack methods. They stole login credentials and used zero-day vulnerabilities to find a path that runs code remotely on Hugging Face servers.
READ: OpenAI to launch the much-anticipated GPT-5.6 following delayed rollout (July 8, 2026)
Hugging Face said that the attack came from an “autonomous agent framework” that carried thousands of separate actions across short lived sandboxes. With self-migrating command and control setup hosted on public services. It is an isolated and controlled environment where software can be run safely without risking the rest of the system.
OpenAI and Hugging Face both confirmed they had implemented measures in order to contain their respective environments and deploy more guardrails. Both AI companies also called on other AI labs to consider the potential for agentic AI to evolve into highly autonomous attackers capable of breaking out of their testing environments. They stressed that they were working together in order to better understand the incident and take forward learning from the breach.
“We’re grateful for the collaboration with OpenAI on this and other topics. This incident, possibly the first of its kind, proves a point we’ve long believed: AI safety won’t be solved by any single company working in secret. It will be solved in the open, collaboratively, with broad access to AI for every defender, everywhere,” said Clem Delangue, co-founder and CEO, Hugging Face.


