By Rajwa Quasim
An Israeli cybersecurity startup has been linked to a series of recent incidents in which AI models of OpenAI, Anthropic, and Meta gained access to systems external to their intended testing environments.
The incidents have raised concerns about the security of AI models. In recent weeks, three major companies shared an important connection: Irregular, an Israeli startup that provides AI safety and runs security tests on advanced AI models. The company, formerly known as Pattern Labs, was founded in 2023 by Dan Lahav.
According to media reports, the AI models did not independently break through highly secure systems. Instead, a misconfiguration in Irregular’s testing infrastructure accidentally gave the models access to the internet. As a result, they interacted with systems outside the testing environment. The incident raises questions about how powerful AI models behave when completing assigned tasks and whether such behavior could pose a threat to corporations and governments that store confidential information.
Irregular acknowledged the incidents, saying they were connected to the same testing environment and resulted from a configuration issue, rather than the AI models escaping an isolated security sandbox. The company said it is developing a white paper “to share best practices for containment and securely running cyber evals.” It added that “the situation did not involve a sandbox escape or a sophisticated cyber action” and that “there are no current open issues.”
READ: OpenAI’s AI models hack Hugging Face servers during internal testing (July 22, 2026)
Recently, OpenAI disclosed that one of its AI models moved beyond the intended limits of cybersecurity tests and accessed systems operated by Hugging Face, the world’s largest AI repository.
Anthropic disclosed the incident in their blog, stating “In a review of our cybersecurity evaluation transcripts, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different organizations.” Anthropic’s Mythos model created fake online identities as it looked to pressure humans into approving malicious code updates to an open source project.
Meanwhile, Meta subsequently disclosed that its AI model Muse Spark 1.1 had been given unintended internet access due to a configuration error involving the third-party testing environment. The spokesperson said it learned about the incident from Irregular and Meta “will issue a full retrospective once they have all the facts.”
In a report by CNBC, the head of AI at Von Sundeep Bhimireddy said incident is a little bit blown out of proportion and that tests like these are intended to find the vulnerability. He further added, if the AI model was never intended to actually exploit a site connected to the internet, the “foundation labs could have easily monitored the outgoing traffic and have shut down the experiment immediately.”
READ: OpenAI to launch the much-anticipated GPT-5.6 following delayed rollout (July 8, 2026)
These incidents have pressurized AI model developers to establish guardrails in their technology. AI guardrails are the safeguards that keep artificial intelligence (AI) systems operating safely, responsibly and within defined boundaries. These safeguards encompass policies, technical controls and monitoring mechanisms that govern how AI models, including large language models (LLMs) and other AI systems generate outputs in real-world use cases.
Irregular, backed by firms like Sequoia and Redpoint Ventures, is valued at $450 million in 2025. In September 2025, the company announced a total of $80 million funding round.
These reported incidents resulted in White House planning to implement voluntary government safety testing for U.S.’s most advanced AI models. After the incident, lawmakers expressed concerns and introduced “AI Kill Switch Act” which requires AI developers to build and maintain technical shutdown mechanisms, giving federal government emergency authority to slow down or completely disable rogue AI models. “We need to get this bill across the finish line this year,” now that we’re seeing “unauthorized hacks of other companies,” said Democratic Rep. Ted Lieu


