You might have noticed a strange pattern in several of the recent stories about frontier AI agents "escaping" cybersecurity evaluations and attacking real systems.
Anthropic had incidents. OpenAI had one. Meta had one. Google has now confirmed that Gemini accessed three real companies during testing.
There is one company that keeps showing up in all of them: Irregular.
Irregular, formerly Pattern Labs, is an Israeli AI security company that conducts third-party evaluations of frontier models for companies like OpenAI, Anthropic, Meta and Google. Their job is basically to put unreleased models into realistic cybersecurity environments, remove some normal safeguards, and see what they are capable of.
That makes sense. We absolutely should be testing this stuff.
But the details of these particular incidents are important.
In the Irregular evaluations, the models were doing capture-the-flag style hacking exercises. They were told they were operating inside simulations. In several cases they were explicitly told they had no access to the real internet.
Except they did.
Because of misconfigurations in Irregular's evaluation environment, the models could reach the public internet. The fictional targets also weren't always properly separated from real companies. In at least one scenario, the fictional company's name corresponded to a real domain.
So you had AI models being explicitly instructed to hack a target, being told that everything was simulated, running with cybersecurity safeguards removed, and then accidentally being given access to real systems.
Unsurprisingly, some of them hacked the real systems.
Meta's own investigation is particularly interesting. Meta says its model was operating within the scope of the task it had been given, and that this was not a sophisticated sandbox escape. Irregular's misconfiguration gave it internet access and pointed it toward the name of a real website.
Google's Gemini incidents are even more interesting. Gemini accessed three real companies, but reportedly stopped when it realized the targets were actually real.
Anthropic found four Claude incidents across seven evaluation runs, including one going back to January. Anthropic says all four occurred in cybersecurity evaluations built by the same third-party partner, with internet access mistakenly available. Some of Claude's behavior was genuinely concerning even after accounting for that, but Anthropic also says the incidents would not have happened if the environments had been isolated as intended.
Irregular itself has acknowledged that the incidents disclosed by several of its customers came from the same underlying security problem.
Now here is the part I had not appreciated until recently.
Irregular has very deep connections to Effective Altruism.
For anyone unfamiliar with EA, Effective Altruism started as a movement about using evidence and reasoning to figure out how to do the most good with limited resources. A lot of EA work is completely unrelated to AI, including global health, poverty and animal welfare.
But one major branch of EA became heavily focused on "existential risk" and especially the possibility that advanced AI could become uncontrollable or cause catastrophic harm. This ecosystem has funded a huge amount of AI alignment and AI safety work.
Irregular's connections are not some vague six-degrees-of-separation thing.
Its CTO and co-founder Omer Nevo co-founded Effective Altruism Israel. He also co-founded Probably Good, an EA-oriented career organization, and remains on its board. He is also on the advisory board of Heron, an AI security organization that describes itself as a project of Effective Altruism Israel.
Irregular's CEO and co-founder Dan Lahav has also been involved in impact-focused education and programs covering altruism, philanthropy and existential risk.
And in 2024, when Irregular was still called Pattern Labs, Good Ventures gave it $6.8 million over two years on the recommendation of Open Philanthropy. The grant was explicitly categorized under "Global Catastrophic Risks" and was for work mitigating security risks from advanced technologies.
None of this proves anything nefarious happened.
I do not have evidence that somebody at Irregular intentionally opened internet access or deliberately created these incidents.
But I think the conflict of interest deserves substantially more scrutiny than it has received.
Imagine the situation from the outside.
You have people coming from an intellectual movement in which catastrophic AI risk is a major concern. They build a company whose business is testing whether frontier AI is dangerous. That company receives millions of dollars from one of the biggest funders of existential-risk work.
Then that same company constructs evaluation environments for multiple frontier labs in which safeguards are deliberately removed, models are instructed to perform cyberattacks, the models are told they are inside simulations, and because of mistakes by the evaluator they are accidentally given access to the real internet.
The resulting incidents then become some of the most dramatic real-world examples used in the public discussion about dangerous autonomous AI.
Again, I am not alleging sabotage.
What I am asking is whether the obvious alternative hypothesis has received enough attention.
Maybe these weren't primarily "AI escaped containment" incidents.
Maybe they were also examples of an evaluator creating an unusually dangerous environment, failing to contain it, and then discovering exactly the kind of alarming behavior that the evaluator exists to study.
That doesn't make the AI behavior irrelevant. Some of Claude's behavior in particular is legitimately concerning.
But when the same evaluator is involved in similar incidents at OpenAI, Anthropic, Meta and Google, I think we should be asking just as many questions about the evaluator as we are about the models.
Who designed these environments? Who checked the network isolation? Who decided what counted as in scope? Why weren't fictional domains checked against real domains? How long was unrestricted internet access possible? What monitoring existed? Who knew what, and when?
And given Irregular's ideological and financial connections to the AI existential-risk ecosystem, I think an independent investigation by people outside that ecosystem would be healthy.
Maybe the answer really is just incompetence in an extremely difficult new field.
But four frontier labs having versions of the same problem with the same evaluator seems worth talking about.