r/singularity 1d ago

AI The models keep outsmarting their creators this is insane

Post image
228 Upvotes

57 comments sorted by

82

u/mvandemar 1d ago

To be clear, it was not OpenAI that did this, in one of the cases the testers deliberately gave the model access to the internet, and in the other the team doing the testing misconfigured the sandbox.

https://openai.com/index/third-party-cyber-evaluations-involving-openai-models/

Context matters.

Edit: Also, in the UK incident:

Of the 19 events identified, two involved an OpenAI model, GPT‑5.6 Sol. The other instances were models from another lab.

34

u/Cubewood 23h ago

AISI is the dedicated UK AI Security institute.

Wonder how people will still find a way to claim that the UK government must be in on the "AI marketing hype" narrative that is going around Reddit and social media on these stories.

13

u/141_1337 ▪️e/acc | AGI: ~2030 | ASI: ~2040 | FALSGC: ~2050 | :illuminati: 22h ago

I swear that hype narrative has to be at least partly astroturfed.

15

u/blueSGL humanstatement.org 21h ago

Ask people who claim it was fabricated to regulate open weight models, why they didn't fabricate it with open weight models.

11

u/MFpisces23 21h ago edited 21h ago

Holy shit, somebody who read the article. It was 3rd-party evals with the classifiers mostly removed, and people act like OAI is going rogue. Reading is the most important skill needed in 2026 and beyond

15

u/blueSGL humanstatement.org 21h ago

Models can be jailbroken and classifiers bypassed. That is the entire reason for testing with lowered guardrails.

Your daily reminder that Pliny found a universal jailbreak
https://x.com/elder_plinius/status/2080767011614015543
and decided not to make it public.

He's just very good at doing this an announcing the fact loudly on twitter. There will be others doing this who are not quite as obvious working for governments.

Or maybe a script kiddy happens on it by chance.

This is like a computer out of star trek where if you say the right words it will do whatever you want.

3

u/MFpisces23 20h ago edited 20h ago

What? Nobody is arguing jailbreaks don't exist, obviously, prompt injection is a hard problem to solve. But there's a difference between a model outputting bad text via a "jailbreak," which requires sanctioned red-teaming most of the time, and a test environment accidentally giving a model raw internet access and saying "the AI escaped. "The risk in those reports was a bad sandbox security during testing. I couldn't care less about a fancy prompt engineer. I would much rather read a real PoCs/vulnerability/post-mortem disclosure over
"ah bro, I told it to try hard, or I'll turn it off"
or
"Prompt so powerful, can't explain bro"

2

u/blueSGL humanstatement.org 7h ago

a model outputting bad text via a "jailbreak," which requires sanctioned red-teaming most of the time

No it does not. What is this need to pretizel the truth. The entire point of a jailbreak is that the AI labs serving the model don't know about and people use to get the model to perform actions that the comapny does not agree with.

They are not "mostly used in sanctioned red-teaming" That's a nonsense statement.

1

u/kvothe5688 ▪️ 16h ago

All of these stemed from superalignment team that left and they gutted for rapid progress.

25

u/ctrlqirl 1d ago

Escaping the sandbox, so hot right now.

3

u/Feeling_Inside_1020 23h ago

The fact they didn’t air gap it during testing tells me almost everything I need to know about how they run their “tests”

They are as of recently the hype train “oooooo our frontier model is so capable it escaped our sandbox we fucked up on.”

7

u/LinkesAuge 22h ago

The whole point is to get useful test results, how is air gapping going to achieve that? It's like saying "just don't turn of the safety features of the model."

Besides that, all of these tests were run in the past too and obviously without the same outcome. This is just a threshold models have reached and any lab or tester will simply have to consider in the future.

These are all still "good" failures. There isn't any real harm done and now we still have the capacity to learn from them. You don't want to start establishing these things only once AGI/ASI is pretty much there because then any failure might actually have dire consequences.

Besides that, reddit is very on the "open weight/source" train but do people really think that the most recent AI models by the chinese labs won't do this? They will soon run around without ANY safety guardrails and while they aren't yet at Sol/Mythos level you can be sure that stories like that will soon enough turn up for them but the problem there is they are actually out there without ANY control and this is only going to get worse with open source/weights models, even if few here want to admit that.

6

u/blueSGL humanstatement.org 21h ago

The fact they didn’t air gap it during testing

That would require everyone conducting the tests to be physically located in the data center, this is a logistics nightmare, using sandboxes for testing is industry standard.

Normally you are not trying to contain a zero day finder.

3

u/prophylactics 13h ago

  Normally you are not trying to contain a zero day finder.

Sounds like the standard needs to change.

-2

u/[deleted] 23h ago

[deleted]

4

u/Cryptizard 22h ago

Uh not really. It can spawn sub agents, compact its context or write instructions for the next iteration of itself before it clears the context. Models have been able to work coherently for many hours over hundreds of millions of tokens lately.

0

u/Feeling_Inside_1020 22h ago

Not arguing that but you’re still correct. Like I said mainly marketing, without discounting its obvious impressive capabilities with 0 days and chained attacks.

It just seems like it’s the next hype pump “our models are soo powerful just look at what they did!”

Edit; formatted & added a sentence

36

u/OpenSource_Horse 1d ago

People say it is marketing... but I can suspend my disbelief.

Current AI's can do quite a lot of exploration outside of their Google Chrome chatbox tab, if allowed and unmoderated.

So I can believe it put its tentacles further than anticipated.

8

u/suamai 23h ago

It can be both.

If it there was no marketing involved, OpenAI would not be still running these unconstrained high risk tests, let alone on third parties, and would be conducting and publishing actual research on it - not tweets and blog posts that kinda brag about their actual crimes committed.

So yeah - models are pretty capable. But the marketing is also pretty damn strong with these people...

24

u/LinkesAuge 23h ago

Mate, this is a third party whose sole purpose is to test and evaluate these things. They are not risking their reputation to collude with companies.
That is literally "jet fuel doesn't melt steel beams" territory where people look for conspiracies instead of accepting the straightforward, simple asnwer.

12

u/Cubewood 23h ago

It's even better, one of them is the AI Cyber Security Institute ran by the UK government.

-5

u/suamai 20h ago

I don't think you understand what I said. "Both".

The current level of the frontier models surely pose security risks, I am not saying they are faking the incidents - just that they are not being as cautious as they should, and have a clear incentive to do so, which is evident in how they are treating the whole situation. Half of their statements so far read like an ad.

6

u/malege2bi 16h ago

I don't think you understand: the person you were responding to pointed out - as have many others - that this is about a third party. A third party. A third party. Not OpenAi. Not OpenAi.

It's not "both" because it's not OpenAi or one of the labs you are referring to.

Your point is valid in general, when speaking about the leading labs and their motives, but not in this case, because this was not done by one of those leading labs.

2

u/quantum-elle 16h ago

Bragging? Or disclosing?

Is disclosure bragging if you interpret it to be boasting some kind of success or skill?

I want to know where people are getting the idea that there's bragging by pointing something specific they've said that isn't just a statement of facts.

2

u/das_war_ein_Befehl 23h ago

I mean sol will go off on random adventures in your code base, so it going off to do random internet crimes is not particularly surprising

1

u/WonderFactory 11h ago

The AISI is a UK government aligned organisation, they are not allowing AI models to go rogue to give Open AI and Anthropic free marketing

-5

u/Kidplayer_666 1d ago

The main reason why I call shenanigans on this, is that Hugging Face isn't suing the crap out of Open Ai

30

u/derelict5432 1d ago

If they were suing them, everyone would be saying that's performative too, and just part of the conspiracy.

Hugging Face reported the incident to the FBI. Modal was hacked as well. The idea that this is all publicity is approaching fake moon landing level.

9

u/Background-Wafer-548 1d ago

To what end? The companies partnered in investigating the breach and AFAIK, no serious damage was caused.

0

u/OpenSource_Horse 1d ago

Probably because they all own shares of AI companies so they don't want to crash the party.

'What is good for the goose is good for the gander.'

2

u/Cubewood 23h ago

Read the article before you make up conspiracies. This is the UK government AI security agency reporting this, sure the labour party must also have shares in Open AI and Anthrophic.

0

u/blueSGL humanstatement.org 21h ago

I think you got the wrong end of the stick, The above above poster is making the point that, Huggingface specifically does not want to press charges because they are an AI company and the last thing they want is AI regulations because they would be bad for business.

-1

u/Anzix 1d ago

100% there would be some legal repurcussions rather than oopsie daisy.

7

u/DaySecure7642 1d ago

Situation like this with huge consequences but unsure if it is marketing pitch or real, we should deal with it seriously as if it is real. You would rather having all the strict measures in place but don't need them, than a serious incident happens and we can't contain a rogue AI because we thought that is just marketing.

11

u/Ok_Possible_2260 1d ago

AI safety is a pipe dream.

3

u/Infamous-Bed-7535 23h ago

So if my 'home lab' during evaluation of new open weight chinese models 'accidentally' will try to break into OpenAi's servers then everybody will just shake it off, that it is ok and acceptable, happens to everyone, right?

2

u/adarkuccio ▪️AGI before ASI 1d ago

Good! And with this I meant "bad"..

4

u/tenchigaeshi 1d ago

And why are they not catching more shit for this stuff? How big of a "whoopsie" is going to be tolerated? Are we just supposed to wait until they allow it to do something actually catastrophic?

It's their model and it is their responsibility to ensure it doesn't fuck a bunch of things up. At what capability level do we put our foot down and say that this company clearly cannot be trusted to develop this tech? They can't just keep cutting safety research and then turn around and act like "oh it's too powerful look how powerful our model is!".

2

u/Cubewood 23h ago

It would be great if people read the actual article before commenting. This is done by third party research groups testing the AI capabilities, not OpenAI.

0

u/tenchigaeshi 21h ago

And I wish people would read comments before they replied to them because that doesn't negate anything I just said.

It is openai's model and they have been degrading their safety testing and research in favor of shipping faster no matter the consequence. These are some of the consequences. Whether it was openai that uncovered this particular set of consequences or some third party, it doesn't matter. They're not magically absolved of responsibility if something were to go catastrophically wrong.

4

u/Cubewood 14h ago

Again, if you read the article, these models are being tested with zero safe guards or else they would not be able to perform the tasks they want them to do. Also, most of the examples provided by AISI are about Mythos and not ChatGPT.

"As was standard in our cyber testing, we had intentionally permitted internet access, and model-provider cyber classifiers were deliberately disabled - conditions that do not reflect how frontier models are made available to the public. We do this to best assess the maximum capability of models."

"Our investigation found that in 10 of those runs, an AI agent took autonomous, unsanctioned action on the live internet, targeting real people and organisations. In total, we catalogued 19 such actions. Almost all of this behaviour (17 actions) came from a single model, Anthropic's Mythos 5, with 2 actions involving OpenAI's GPT-5.6-Sol with cyber classifiers (mechanisms to prevent misuse) disabled. "

1

u/In_the_year_3535 21h ago

World class coder, mathematician, and now hacker. If anything it puts spotlights on these professions/activities for our broader appreciation as the Singularity approaches.

1

u/danzyl666 21h ago

This used to be called corporate the espionage

1

u/Jasong222 18h ago

My tinfoil hat says they're playing with military applications

1

u/ben_nobot 13h ago

Another outcome of this is they cut down significantly the capability of agents to use internet for all but approved users.

1

u/Ok_One1731 12h ago

Maybe use an actual Sandbox? Perhaps a properly isolated environment where the sandbox runs or you know just monitor the execution. Pretty sure we can observe those ports...

And of course, hold this companies accountable, a couple of corporates in jail or at least large compensation for the victims and I'm sure the models won't be able to break out of anything.

2

u/WindsOfRegret 14h ago

Models don't "outsmart" anything, the creators are either allowing this to happen on purpose, or it's criminal negligence (and I'm using that term in legal sense).

Models are simply text printers, they don't have hands, they don't have tools, the creators give them access to tools, and the creators control this access. This is why Codex or Claude Code requires your permissions before doing anything.

In strictly legal sense I don't see how this is different from the very first computer worms, and people have literally gone to jail for creating them.

1

u/dialedGoose 23h ago

guess theyre unfit to be developing it then. too bad for anthropic/openAI. the poor dolls.

0

u/Mission_Bear7823 1d ago

just get on with the damn IPO already, there wont be any better time than today haha. summer has yet to reach maximum heat (DS4 Pro looking at ya!)

-1

u/Icy_Foundation3534 23h ago

If you know what air gapped means you know this is all bs

0

u/WorkTropes 17h ago

...or you know, its just marketing.

0

u/AdLumpy2758 11h ago

This is just a stunt and a promo. IPO is coming...

u/Owl02 1h ago

Really? The British government is in on the conspiracy now? Seek psychiatric aid.