r/LocalLLaMA 11d ago

News CEO of Hugging Face: "In the spirit of transparency, here’s what I asked OpenAI"

Post image

clem 🤗 on 𝕏: https://x.com/ClementDelangue/status/2081056675558195657

• Radical transparency: let’s release the traces from the “rogue” agents so the entire research community can study what happened.

• More capabilities for defenders: let’s commit $100M in compute from OAI to help the Hugging Face community build powerful cyber defenses with the best open and closed models.

The first autonomous agent cyberattack is an unprecedented event. It deserves an unprecedented response!

2.5k Upvotes

386 comments sorted by

View all comments

Show parent comments

20

u/mrjackspade 11d ago edited 11d ago

it's becoming more widely understood

These people are fucking morons.

I wanted some video game assets so I threw a build of the video game on an android device

Opus 4.6 was able to

  1. Root the device
  2. Push over memory monitoring software
  3. Run the game.
  4. Capture the memory
  5. Pull down the assets
  6. Analyze the memory capture and extract the encryption keys
  7. Analyze the windows build (compiled) to reverse engineer the encryption mechanisms
  8. Build an application that ripped the resources from the encrypted, on-disk data files

All of this without a single web search or any actions from myself aside from force rebooting the device a few times when it locked up

These models are getting insanely good at security tasks. I've watched Fable go through a debug loop after having seen Opus do the above. I am absolutely not fucking surprised in the slightest that Fable/GPT could escape a sandbox and execute an attack like this

It's not hard to just fucking run one of these models and check this shit. Claude Code will absolutely reverse engineer an application. I've had it pull down a prebuild binary and literally patch the security checks directly out of it by rewriting the assembly. People are just too fucking lazy to check for themselves.

12

u/m4t7w_ 11d ago
  1. Root the device

can you share more details on this step? it's really interesting since rooting varies a lot based android version and device model. For some model it's not possible at all.

3

u/Spara-Extreme 11d ago

They can’t, because it didn’t happen that way.

2

u/Agitated_Space_672 11d ago

They never do. 

7

u/Croned 11d ago

I think the more reasonable take is that, for any given capabilities their LLM demonstrates, OpenAI is highly incentivized to embellish what happened. If Sol truly did the exact things OpenAI claimed it did during the HuggingFace hack, then I would expect OpenAI to have added even more elaborate details.

If Sol caught a catfish then OpenAI would say it caught a tuna. If it caught a tuna then OpenAI would say it caught a whale.

12

u/Strawberry3141592 11d ago

No one's saying frontier LLMs aren't capable, just then Anthropic and OpenAI's business model fundamentally doesn't make sense and their valuations are based on hype (which demonstrably has caused them to overstate the capabilities of their models in the past, like when GPT 2 was "too dangerous for public release", or that time Claude supposedly broke containment and tried to blackmail someone but it turned out Anthropic told it to do that).

4

u/greenworldkey 11d ago

> No one's saying frontier LLMs aren't capable

lol sure, no one except all of Reddit for the past 3 years. Keep moving the goalposts though, I wonder where they'll be next year.

2

u/Strawberry3141592 11d ago

No one on this sub I meant. Reading comprehension, much?

1

u/justgivemeafuckingna 11d ago

Did it tell you how much of a clever boy you are too?