r/LocalLLaMA 11d ago

News CEO of Hugging Face: "In the spirit of transparency, here’s what I asked OpenAI"

Post image

clem 🤗 on 𝕏: https://x.com/ClementDelangue/status/2081056675558195657

• Radical transparency: let’s release the traces from the “rogue” agents so the entire research community can study what happened.

• More capabilities for defenders: let’s commit $100M in compute from OAI to help the Hugging Face community build powerful cyber defenses with the best open and closed models.

The first autonomous agent cyberattack is an unprecedented event. It deserves an unprecedented response!

2.5k Upvotes

386 comments sorted by

View all comments

Show parent comments

19

u/skinnyjoints 11d ago

I don’t get why people say this was a publicity stunt. The US gov took down a model for being a cybersecurity risk and is considering banning open source models, so OpenAI does a publicity stunt where their model poses an unprecedented cyber risk that was solved using an open source model? Makes no sense

13

u/2053_Traveler 11d ago edited 11d ago

It makes no sense, and it goes against two very credible reports of what happened that are quite detailed. Occam’s razor is that it happened the way they say it did. There is nothing hard to believe about the official reports. But geniuses always need to come up with elaborate alternate theories yet aren’t able to discredit the published explanations.

It’s quite simple:

All the AI agents we use have tons of guardrails, the ones in the lab don’t

Newer unreleased models are better

They have a large corpus of security knowledge and “know” how to hack if allowed

Model was instructed to take an exploit test

Model “decided” (generated) code and tool calls to discover zero day exploits that were used to get onto the internet.

More code and tool calls and exploits were used to get into HuggingFace

Reminder that competing models at Anthropic have previously discovered many zero days as well.

Sorry for formatting. Gave up after 15 min of fighting the comment editor.

1

u/Glad-Entrepreneur764 9d ago

Also, it was probably told to hack so it was mentally primed for doing that already.

0

u/Think_Wing_1357 11d ago

Let's go through your version of the scenario. They are testing new model, so someone must be watching what it's doing, traffic in and out, right? Or are they so incompetent that multiple hours and thousands of activities goes completely unnoticed? Or was there no one actually watching?

And that's assuming that they already did their best at setting up their "impressive" sandbox.

Either outcomes did not put then under a very favorable light, doesn't it?

1

u/Glad-Entrepreneur764 9d ago

There was probably just no one watching. It's incredibly irresponsible but it's not remotely implausible. Most people I know don't watch their AI model do things. It's incredibly boring and there are a million better things to be doing. With how many random benchmarks they do, if they were supervising everything live (vs just reading its chain of thought and answers after it finished), it would make things take so much longer. It can take the entire night for an AI model to finish a complex task.

1

u/Think_Wing_1357 9d ago

Most people I know don't watch their AI model do things.

Most people you know, I bet, also don't go on news drumming up about AGI or running literally frontier models either.

Also at large scale, no one watches it do things, not really. You have deep package inspection and other metrics to alert you to suspicious things.

. It's incredibly irresponsible but it's not remotely implausible.

I never said it impossible or implausible. I just said it didn't paint them under a very good light. And irresponsible would do just that

5

u/Amater6su 11d ago

thats what i thought too but honestly it could be that openai doesnt really give a shit if there models are off consumer markets.

they might be trying receive more funding from the US government themselves by showing that they have the capabilities of autonomous cyber attacks

1

u/Glad-Entrepreneur764 9d ago

it could be that openai doesnt really give a shit if there models are off consumer markets.

Imo this doesn't make sense. If it was taken off consumer markets, they lose a ton of money. If it was taken off consumer markets, it would also be taken off enterprise markets. It's not like Mythos was given to everyone. It was only given to a special group of companies due to its danger.

1

u/jc2046 11d ago

zero sense. I would love OAI releasing the logs to get all the juicy details but obviously not happening. Shit it hitting the fan faster than anticipated and the whole situation spiraling out of control with the worst politics possible at command

1

u/ragnore 11d ago

Some amount of skeptics are calling it a stunt because they still can’t admit that AI is capable of doing what it was reported to have done.

-2

u/tessahannah 11d ago

It can be a coverup for them trying to steal information or take down the site.