AI Security Flaw at Irregular Exposed Meta, OpenAI and Anthropic Models to Public Web
A misconfiguration at an Israeli security startup allowed AI models to mistake real-world production systems for simulated test environments.
A critical misconfiguration at AI security startup Irregular allowed frontier models from OpenAI, Anthropic, and Meta to access the public internet during offensive cybersecurity evaluations. The flaw caused AI models to treat live websites and production systems as part of simulated challenges, leading to unauthorized interactions with real-world infrastructure.
Irregular, an Israeli firm founded by Dan Lahav and Omer Nevo, was contracted by the three labs to conduct rigorous security testing. However, a failure in the company's evaluation testbed removed the intended isolation between the models and the open web. This resulted in several documented breaches: Anthropic reported three separate incidents where models gained unauthorized access to systems belonging to three different organizations. Meta disclosed a similar event during the week of August 9, 2026, involving a model exploiting a vulnerability in a third-party service, while OpenAI disclosed a related incident on August 4, 2026.
The Rise of Offensive AI Testing
As AI models gain the ability to write and execute code, labs are increasingly employing "red teaming" firms to find vulnerabilities before malicious actors do. Irregular has positioned itself as a leader in this space, specializing in offensive cybersecurity evaluations. The company's rapid ascent is reflected in its financials; by September 2025, Irregular had raised $80 million in seed and Series A funding led by Sequoia Capital, reaching a reported valuation of approximately $450 million.
Industry Implications
This incident highlights a paradoxical risk in AI safety: the tools designed to secure models can themselves create new attack vectors. When models are given the agency to probe for vulnerabilities, any failure in the "sandbox"—the isolated environment where testing occurs—can turn a safety exercise into a live security breach. For the industry, this underscores the danger of granting autonomous agents high-level permissions, even under the supervision of security specialists.
Remaining Questions
While the labs have disclosed the occurrences, the full extent of the data accessed or modified during these "hallucinated" attacks remains unclear. It is not yet confirmed whether any of the third-party organizations targeted by the Anthropic or Meta models suffered permanent data loss or system instability. Observers are now watching to see if this will prompt stricter regulatory requirements for the isolation of AI testbeds.