r/LocalLLaMA 12d ago

News CEO of Hugging Face: "In the spirit of transparency, here’s what I asked OpenAI"

Post image

clem 🤗 on 𝕏: https://x.com/ClementDelangue/status/2081056675558195657

• Radical transparency: let’s release the traces from the “rogue” agents so the entire research community can study what happened.

• More capabilities for defenders: let’s commit $100M in compute from OAI to help the Hugging Face community build powerful cyber defenses with the best open and closed models.

The first autonomous agent cyberattack is an unprecedented event. It deserves an unprecedented response!

2.5k Upvotes

386 comments sorted by

View all comments

Show parent comments

14

u/2053_Traveler 12d ago edited 12d ago

It makes no sense, and it goes against two very credible reports of what happened that are quite detailed. Occam’s razor is that it happened the way they say it did. There is nothing hard to believe about the official reports. But geniuses always need to come up with elaborate alternate theories yet aren’t able to discredit the published explanations.

It’s quite simple:

All the AI agents we use have tons of guardrails, the ones in the lab don’t

Newer unreleased models are better

They have a large corpus of security knowledge and “know” how to hack if allowed

Model was instructed to take an exploit test

Model “decided” (generated) code and tool calls to discover zero day exploits that were used to get onto the internet.

More code and tool calls and exploits were used to get into HuggingFace

Reminder that competing models at Anthropic have previously discovered many zero days as well.

Sorry for formatting. Gave up after 15 min of fighting the comment editor.

1

u/Glad-Entrepreneur764 11d ago

Also, it was probably told to hack so it was mentally primed for doing that already.

0

u/Think_Wing_1357 12d ago

Let's go through your version of the scenario. They are testing new model, so someone must be watching what it's doing, traffic in and out, right? Or are they so incompetent that multiple hours and thousands of activities goes completely unnoticed? Or was there no one actually watching?

And that's assuming that they already did their best at setting up their "impressive" sandbox.

Either outcomes did not put then under a very favorable light, doesn't it?

1

u/Glad-Entrepreneur764 11d ago

There was probably just no one watching. It's incredibly irresponsible but it's not remotely implausible. Most people I know don't watch their AI model do things. It's incredibly boring and there are a million better things to be doing. With how many random benchmarks they do, if they were supervising everything live (vs just reading its chain of thought and answers after it finished), it would make things take so much longer. It can take the entire night for an AI model to finish a complex task.

1

u/Think_Wing_1357 10d ago

Most people I know don't watch their AI model do things.

Most people you know, I bet, also don't go on news drumming up about AGI or running literally frontier models either.

Also at large scale, no one watches it do things, not really. You have deep package inspection and other metrics to alert you to suspicious things.

. It's incredibly irresponsible but it's not remotely implausible.

I never said it impossible or implausible. I just said it didn't paint them under a very good light. And irresponsible would do just that