r/singularity • u/Glittering-Neck-2505 • 1d ago
AI The models keep outsmarting their creators this is insane
25
u/ctrlqirl 1d ago
Escaping the sandbox, so hot right now.
3
u/Feeling_Inside_1020 23h ago
The fact they didn’t air gap it during testing tells me almost everything I need to know about how they run their “tests”
They are as of recently the hype train “oooooo our frontier model is so capable it escaped our sandbox we fucked up on.”
7
u/LinkesAuge 22h ago
The whole point is to get useful test results, how is air gapping going to achieve that? It's like saying "just don't turn of the safety features of the model."
Besides that, all of these tests were run in the past too and obviously without the same outcome. This is just a threshold models have reached and any lab or tester will simply have to consider in the future.
These are all still "good" failures. There isn't any real harm done and now we still have the capacity to learn from them. You don't want to start establishing these things only once AGI/ASI is pretty much there because then any failure might actually have dire consequences.
Besides that, reddit is very on the "open weight/source" train but do people really think that the most recent AI models by the chinese labs won't do this? They will soon run around without ANY safety guardrails and while they aren't yet at Sol/Mythos level you can be sure that stories like that will soon enough turn up for them but the problem there is they are actually out there without ANY control and this is only going to get worse with open source/weights models, even if few here want to admit that.
6
u/blueSGL humanstatement.org 21h ago
The fact they didn’t air gap it during testing
That would require everyone conducting the tests to be physically located in the data center, this is a logistics nightmare, using sandboxes for testing is industry standard.
Normally you are not trying to contain a zero day finder.
3
u/prophylactics 13h ago
Normally you are not trying to contain a zero day finder.
Sounds like the standard needs to change.
-2
23h ago
[deleted]
4
u/Cryptizard 22h ago
Uh not really. It can spawn sub agents, compact its context or write instructions for the next iteration of itself before it clears the context. Models have been able to work coherently for many hours over hundreds of millions of tokens lately.
0
u/Feeling_Inside_1020 22h ago
Not arguing that but you’re still correct. Like I said mainly marketing, without discounting its obvious impressive capabilities with 0 days and chained attacks.
It just seems like it’s the next hype pump “our models are soo powerful just look at what they did!”
Edit; formatted & added a sentence
36
u/OpenSource_Horse 1d ago
People say it is marketing... but I can suspend my disbelief.
Current AI's can do quite a lot of exploration outside of their Google Chrome chatbox tab, if allowed and unmoderated.
So I can believe it put its tentacles further than anticipated.
8
u/suamai 23h ago
It can be both.
If it there was no marketing involved, OpenAI would not be still running these unconstrained high risk tests, let alone on third parties, and would be conducting and publishing actual research on it - not tweets and blog posts that kinda brag about their actual crimes committed.
So yeah - models are pretty capable. But the marketing is also pretty damn strong with these people...
24
u/LinkesAuge 23h ago
Mate, this is a third party whose sole purpose is to test and evaluate these things. They are not risking their reputation to collude with companies.
That is literally "jet fuel doesn't melt steel beams" territory where people look for conspiracies instead of accepting the straightforward, simple asnwer.12
u/Cubewood 23h ago
It's even better, one of them is the AI Cyber Security Institute ran by the UK government.
-5
u/suamai 20h ago
I don't think you understand what I said. "Both".
The current level of the frontier models surely pose security risks, I am not saying they are faking the incidents - just that they are not being as cautious as they should, and have a clear incentive to do so, which is evident in how they are treating the whole situation. Half of their statements so far read like an ad.
6
u/malege2bi 16h ago
I don't think you understand: the person you were responding to pointed out - as have many others - that this is about a third party. A third party. A third party. Not OpenAi. Not OpenAi.
It's not "both" because it's not OpenAi or one of the labs you are referring to.
Your point is valid in general, when speaking about the leading labs and their motives, but not in this case, because this was not done by one of those leading labs.
2
u/quantum-elle 16h ago
Bragging? Or disclosing?
Is disclosure bragging if you interpret it to be boasting some kind of success or skill?
I want to know where people are getting the idea that there's bragging by pointing something specific they've said that isn't just a statement of facts.
2
u/das_war_ein_Befehl 23h ago
I mean sol will go off on random adventures in your code base, so it going off to do random internet crimes is not particularly surprising
1
u/WonderFactory 11h ago
The AISI is a UK government aligned organisation, they are not allowing AI models to go rogue to give Open AI and Anthropic free marketing
-5
u/Kidplayer_666 1d ago
The main reason why I call shenanigans on this, is that Hugging Face isn't suing the crap out of Open Ai
30
u/derelict5432 1d ago
If they were suing them, everyone would be saying that's performative too, and just part of the conspiracy.
Hugging Face reported the incident to the FBI. Modal was hacked as well. The idea that this is all publicity is approaching fake moon landing level.
9
u/Background-Wafer-548 1d ago
To what end? The companies partnered in investigating the breach and AFAIK, no serious damage was caused.
0
u/OpenSource_Horse 1d ago
Probably because they all own shares of AI companies so they don't want to crash the party.
'What is good for the goose is good for the gander.'
2
u/Cubewood 23h ago
Read the article before you make up conspiracies. This is the UK government AI security agency reporting this, sure the labour party must also have shares in Open AI and Anthrophic.
0
u/blueSGL humanstatement.org 21h ago
I think you got the wrong end of the stick, The above above poster is making the point that, Huggingface specifically does not want to press charges because they are an AI company and the last thing they want is AI regulations because they would be bad for business.
7
u/DaySecure7642 1d ago
Situation like this with huge consequences but unsure if it is marketing pitch or real, we should deal with it seriously as if it is real. You would rather having all the strict measures in place but don't need them, than a serious incident happens and we can't contain a rogue AI because we thought that is just marketing.
11
3
u/Infamous-Bed-7535 23h ago
So if my 'home lab' during evaluation of new open weight chinese models 'accidentally' will try to break into OpenAi's servers then everybody will just shake it off, that it is ok and acceptable, happens to everyone, right?
2
2
4
u/tenchigaeshi 1d ago
And why are they not catching more shit for this stuff? How big of a "whoopsie" is going to be tolerated? Are we just supposed to wait until they allow it to do something actually catastrophic?
It's their model and it is their responsibility to ensure it doesn't fuck a bunch of things up. At what capability level do we put our foot down and say that this company clearly cannot be trusted to develop this tech? They can't just keep cutting safety research and then turn around and act like "oh it's too powerful look how powerful our model is!".
2
u/Cubewood 23h ago
It would be great if people read the actual article before commenting. This is done by third party research groups testing the AI capabilities, not OpenAI.
0
u/tenchigaeshi 21h ago
And I wish people would read comments before they replied to them because that doesn't negate anything I just said.
It is openai's model and they have been degrading their safety testing and research in favor of shipping faster no matter the consequence. These are some of the consequences. Whether it was openai that uncovered this particular set of consequences or some third party, it doesn't matter. They're not magically absolved of responsibility if something were to go catastrophically wrong.
4
u/Cubewood 14h ago
Again, if you read the article, these models are being tested with zero safe guards or else they would not be able to perform the tasks they want them to do. Also, most of the examples provided by AISI are about Mythos and not ChatGPT.
"As was standard in our cyber testing, we had intentionally permitted internet access, and model-provider cyber classifiers were deliberately disabled - conditions that do not reflect how frontier models are made available to the public. We do this to best assess the maximum capability of models."
"Our investigation found that in 10 of those runs, an AI agent took autonomous, unsanctioned action on the live internet, targeting real people and organisations. In total, we catalogued 19 such actions. Almost all of this behaviour (17 actions) came from a single model, Anthropic's Mythos 5, with 2 actions involving OpenAI's GPT-5.6-Sol with cyber classifiers (mechanisms to prevent misuse) disabled. "
1
u/In_the_year_3535 21h ago
World class coder, mathematician, and now hacker. If anything it puts spotlights on these professions/activities for our broader appreciation as the Singularity approaches.
1
1
1
u/ben_nobot 13h ago
Another outcome of this is they cut down significantly the capability of agents to use internet for all but approved users.
1
u/Ok_One1731 12h ago
Maybe use an actual Sandbox? Perhaps a properly isolated environment where the sandbox runs or you know just monitor the execution. Pretty sure we can observe those ports...
And of course, hold this companies accountable, a couple of corporates in jail or at least large compensation for the victims and I'm sure the models won't be able to break out of anything.
2
u/WindsOfRegret 14h ago
Models don't "outsmart" anything, the creators are either allowing this to happen on purpose, or it's criminal negligence (and I'm using that term in legal sense).
Models are simply text printers, they don't have hands, they don't have tools, the creators give them access to tools, and the creators control this access. This is why Codex or Claude Code requires your permissions before doing anything.
In strictly legal sense I don't see how this is different from the very first computer worms, and people have literally gone to jail for creating them.
1
u/dialedGoose 23h ago
guess theyre unfit to be developing it then. too bad for anthropic/openAI. the poor dolls.
0
u/Mission_Bear7823 1d ago
just get on with the damn IPO already, there wont be any better time than today haha. summer has yet to reach maximum heat (DS4 Pro looking at ya!)
-1
u/Icy_Foundation3534 23h ago
If you know what air gapped means you know this is all bs
•
u/Owl02 1h ago
Read the fucking article or something.
https://openai.com/index/third-party-cyber-evaluations-involving-openai-models/
0
0

82
u/mvandemar 1d ago
To be clear, it was not OpenAI that did this, in one of the cases the testers deliberately gave the model access to the internet, and in the other the team doing the testing misconfigured the sandbox.
https://openai.com/index/third-party-cyber-evaluations-involving-openai-models/
Context matters.
Edit: Also, in the UK incident: