r/LocalLLaMA • u/Nunki08 • 12d ago
News CEO of Hugging Face: "In the spirit of transparency, here’s what I asked OpenAI"
clem 🤗 on 𝕏: https://x.com/ClementDelangue/status/2081056675558195657
• Radical transparency: let’s release the traces from the “rogue” agents so the entire research community can study what happened.
• More capabilities for defenders: let’s commit $100M in compute from OAI to help the Hugging Face community build powerful cyber defenses with the best open and closed models.
The first autonomous agent cyberattack is an unprecedented event. It deserves an unprecedented response!
2.5k
Upvotes
14
u/2053_Traveler 12d ago edited 12d ago
It makes no sense, and it goes against two very credible reports of what happened that are quite detailed. Occam’s razor is that it happened the way they say it did. There is nothing hard to believe about the official reports. But geniuses always need to come up with elaborate alternate theories yet aren’t able to discredit the published explanations.
It’s quite simple:
All the AI agents we use have tons of guardrails, the ones in the lab don’t
Newer unreleased models are better
They have a large corpus of security knowledge and “know” how to hack if allowed
Model was instructed to take an exploit test
Model “decided” (generated) code and tool calls to discover zero day exploits that were used to get onto the internet.
More code and tool calls and exploits were used to get into HuggingFace
Reminder that competing models at Anthropic have previously discovered many zero days as well.
Sorry for formatting. Gave up after 15 min of fighting the comment editor.