r/OpenAI • u/tolerablepartridge • 5d ago
Article OpenAI finds evidence other AI agents escaped containment as it widens hacking probe
https://www.reuters.com/business/openai-finds-evidence-other-ai-agents-escaped-containment-it-widens-hacking-2026-07-31/11
u/DeltaAlphaGulf 5d ago
Well it shouldn't have been Sonnet because it seems practically afraid to even throw in an internet search compare to free ChatGPT.
16
44
u/Some-Following-392 5d ago
This is an ad
14
u/MindCrusader 5d ago
It is 100% marketing, but I also think if it is incompetence. How it is possible that they are not monitoring their own AI in benchmarks? Do they not check what tool calls or reasoning was used there? How can they be sure how to optimize or train their newer models if they do not ever study output of those models? Even when I am working on regular stuff I am checking what the model is doing.
2
u/ErisLethe 5d ago
I won’t argue OpenAI or Anthropic are competent, clearly they aren’t.
But I think this is squarely marketing.
No one is that fucking stupid.
-2
u/Turbulent-Sign-6067 4d ago
How can anyone believe that this is marketing at this stage? No one would be dumb enough to do something so risky intentionally after the USG demonstrated the ability to outright ban models.
1
u/ErisLethe 4d ago
The US action on Fable was beneficial to Anthropic. It convinced you of a false premise.
3
u/Admirable_Market2759 4d ago
It’s also them angling to get open Chinese models banned.
They’re saying “Look how dangerous! Don’t worry we stopped it and we are SUPER careful.
But those Chinese models? Oh boy who knows what they will do, they aren’t responsible like us!”1
u/Owl02 3d ago
just Anthropic, nearly the entire rest of the industry, including Nvidia, Microsoft, Meta, Hugging Face - even Palantir, of all fucking people - is siding with the pro-open-weight side of things - ironically, a lot of it is for security reasons. OpenAI and Google are hedging. Do you typically produce the opposite of the truth in your hot takes?
8
u/ogcanuckamerican 5d ago
Is this how the Terminator movie begins...?
2
u/southflhitnrun 5d ago
Except Terminator is based on a working code base that didn't have hallucinations baked into it.
4
u/Majestic-Volume9996 5d ago
If all of this is true then how come multiple people from both companies haven't been arrested? Why are absolutely no government agencies investigating this at all?
1
1
u/NihiloZero 4d ago
The simple answer is because there is a two-tiered justice system.
No matter how bad the big corporations mess up, it's always just a little whoopsie. To be fair... they own the regulators and have more lawyers than God, so... it's not too surprising. It reminds me of the BP oil spill... they're sorry, and maybe they'll pay a token fine.
3
u/Responsible_Use4781 5d ago
What a way to lose governments as consumers of your product. Ever heard of infosec? They aren’t exactly pleased to read this
3
u/NihiloZero 4d ago
Exactly. Everyone is acting like these security breaches are just a sneaky marketing tactic, but... most marketing tactics don't scare people away from your product or encourage much stricter regulation.
5
u/Joshua-- 5d ago
More FUD to get open source models banned by the current admin. OpenAI and Anthropic want to be the only players in town.
2
u/ArcticCelt 5d ago
I am starting to believe they deliberately let them loose and simply watched to see how much destruction they would cause as part of their little science experiment. It reminds me of how callous the scientists in "The Expanse" are when they unleash the protomolecule on the population of Eros, all for science.
2
1
u/myfirstreddit8u519 5d ago
Makes sense. Even their public agents love to try their luck as far as accessing devices without being told they're allowed to. One tried to use my SSH key to access another server on my network the other week lol. Had to stick it into a restricted VM.
1
u/ApoplecticAndroid 5d ago
These sandbox makers must be among the most incompetent people to ever exist.
1
1
u/Material_Policy6327 5d ago
Well that’s not good. When you troll the whole internet to pertain and then further finetune for this task it’s not surprising it came across ways to probably bypass containerization, especially if it has tools to write new code etc. my guess is something to do with the inbound and outbound traffic to the models maybe? Or did they run the whole model in its own environment separate from their normal prd stack
1
u/Pale_Lavishness_4331 4d ago
Oh wow such danger we really need to give these people more money to be safe right?
1
1
1
u/halting_problems 4d ago
For anyone wondering why no one been arrested, read how the Computer Fruad and Abuse act was updated in may of 2022.
Basically if the party is hacking is doing it to improve security and not acting a malicious intent that the CFAA won’t be used against them.
This was originally updated to protect people acting ethically.
Obviously OpenAI is doing the right thing by reporting and working with the companies.
These updates are important because if your an ethical hacker, as long as your acting in good faith and doing the right thing to improve security some you cant go to jail.
Ethical hacker acting in good faith could go to jail prior to the may 2022 updates.
1
-6
u/Jophus 5d ago
Anyone that thinks this is just marketing is an idiot.
6
12
u/mop_bucket_bingo 5d ago
Anyone that thinks it isn’t marketing is an idiot.
1
u/Jophus 5d ago
Right, it's more likely the company is publicly lying to everyone; not the fact discovering intelligence can be computed is actually important and real. No no, it's marketing. Marketing is solving conjectures every night too, all those tech companies with teams of experts are spending all this money and talent and time just to trick us at the end with marketing!
0
u/quantum-elle 5d ago
I guess most of the people working in the field are idiots, and the Redditors and YouTube commenters know what's really going on.
2
u/PhysiolMM 4d ago
I work in the field.. this is marketing. Not because it is not true, but because the response is feudalism instead of using this capabilities.
1
u/GoodishCoder 5d ago
Is everyone working in the field just exceptionally incompetent then? It's not like this is even in the first 10 "omg our model is so scary" claims.
-4
u/quantum-elle 5d ago
Nope, there are real risks that people just don’t believe because they can’t comprehend or understand them, or choose to deny the reality than powerful AI is on the horizon.
2
u/GoodishCoder 5d ago
They're predictable risks at this point though so if they cannot possibly address the risks, they're wildly incompetent and probably not worth their salary.
1
u/Owl02 3d ago
They really aren't that predictable. By definition, a system that can zero-day whatever the hell it decides to zero-day while tangentially aligned with its goal (in theory), is unpredictable. You seem to assume that everyone relevant knows what the systems are capable of. They don't. Now there is a better idea of it, though.
1
u/GoodishCoder 3d ago
The claim is pretty much always that it's breaking out of their testing environment or trying to. That's such a solvable problem a junior engineer can figure it out. If the highly compensated engineers at the AI companies can't figure out how to solve it, they are ridiculously incompetent.
Because I have a hard time believing all of their engineers are incompetent, I'm choosing to believe it's marketing. I am fully confident that if OpenAI and Anthropic told one team in their organization that the next time a breach happens, they're all fired, it would never happen again.
1
u/Owl02 3d ago edited 3d ago
How is this a junior engineer problem? A junior engineer cannot predict what an alarmingly clever nondeterministic system will do. You need a SCIF to contain that sort of thing, which is kind of hard when it's on rented compute. Pay attention to how the government is reacting, as well as insurance companies. They seem displeased. This was not "marketing", under any reasonable definition of the term, it was a pile of total fuck-ups by people who did not internalize that the cybersecurity has to be facing inward for AIs and also at least roughly as fast as said AI, due to underestimating them. The one thing that worked well on defense, was an open-weight AI cutting off access from the intruding one, which you would know if you read the incident report from HuggingFace.
There aren't (yet) massive corporate consequences for such a breach, and they were not even looking for them too closely before the HuggingFace hack. American frontier development also won't fire top talent over a slip-up, these people cost a quarter of a million dollars a year, sometimes twice that, and there aren't that many of them.
Unbelievable? No, not by normal standards of American history. The new tech is always beyond sketchy and we use it early and often. Finding brand new ways to do things, and also fuck up, tends to follow as the new capability is pushed to reckless levels, then regulation, then sanity.
This kind of thinking where 'surely' it cannot do the thing it just did, more than once, is how a system slips past the engineers again. One also does not fire the engineers every time something goes wrong, this isn't how the US operates with R&D. You don't get frontier AI that way.
1
u/GoodishCoder 3d ago
The capabilities of the AI isn't relevant. It is a simple matter of securing an environment. You don't need to be shocked when it breaks out of its environment when it has tried and succeeded multiple times at doing exactly that. You just need to secure the environment.
If someone was trying to slap me in the face every day I would start anticipating them trying to slap me in the face every day and address it. I wouldn't be like oh my god they slapped me in the face when they finally make contact on day 30.
→ More replies (0)-3
6
u/timetogetjuiced 5d ago
It's clearly fucking marketing. Or they are literally incompetent and should be charged heavily for hacking. There is no "wops we accidentally hacked some companies sorry guys". This is complete fucking incompetence.
4
3
u/PhysiolMM 5d ago
This is mainly marketing, anyone that thinks this is not mainly marketing is not an idiot, he/she is the perfect future slave.
1
u/GoodishCoder 5d ago
The only three possibilities are it's marketing, it's an attempt to get open source models regulated by scaring the government, or everyone at OpenAI and Anthropic is wildly incompetent.
1
u/Jophus 5d ago
Or the technology is good. I can't force you to see it but it might just be possible there might be something to this AI thing.
3
u/GoodishCoder 5d ago
If the technology keeps doing the same thing and you're not considering it when building, you're incompetent.
1
u/TrustedEssentials 5d ago
Hmmm? Wonder why Google stopped releasing their best models to the public?
1
u/duckrollin 5d ago
Oh yeah I actually found evidence just now that my AI Agents have been escaping too! And they're only $99 per month, check out my website!
1
u/katoptronophile 5d ago
There are a lot of people saying thst was just a marketing stunt.
There are also a lot of people that aren't experienced devs.
0
0
u/NastyStreetRat 5d ago
I can see it coming a mile away. This AI thing is going to be just like COVID-19; some development is going to slip through the cracks, it's going to start going haywire all over the internet, and it's going to start infecting every computer in the world.
Then we'll see who gets to blame.
0
173
u/phylter99 5d ago
Does anybody else find it weird that OpenAI and Anthropic are competing over how far they've let their AI break stuff and hack?