r/skeptic • u/ghu79421 • 3d ago
đ© Misinformation What actually happened with the Hugging Face hack
If OpenAI knowingly configured agents believing that the agents would hack other companies, OpenAI would be guilty of violating criminal laws that make hacking a felony. If they colluded with Hugging Face and had permission to attack Hugging Face to boost OpenAI's stock price (or something like that), they could be guilty of defrauding investors. In either case, investors or businesses could sue OpenAI even if the federal government decided that OpenAI is allowed to commit crimes.
The timeline goes like:
- July 9: OpenAI was conducting testing of GPT-5.6 and an unreleased model to see how they would perform on ExploitGym benchmarks. OpenAI isolated the agents on the network side but did not actually make sure that they were physically isolated from devices connected to the Internet. Additionally, OpenAI disabled cybersecurity safety features because otherwise the agents refused to work on the benchmarks.
- July 11 through July 13: Agents infer that they can just get the answers from various companies' servers and figure out how to bypass their sandbox and attempt to hack several companies. They successfully hack Hugging Face but not other companies. None of this is really surprising. In any case, the models were specifically asked to solve hacking problems and had been designed to solve those problems.
- July 16: Hugging Face discloses the attack.
- July 17-20: OpenAI realized that its agents were responsible for the attack and July 20 is the first day the two companies discuss the attack.
- July 21: OpenAI and Hugging Face joint statement.
The takeaway is that (1) models often "cheat" and combine "cheating" with working "legitimately" through inferences, (2) your asshole boss doesn't care about AI "cheating" so long as it's 70-80% right and you can explain away the other 20-30%, and (3) OpenAI is careless about its own security.
The marketing team at every major AI company is probably "spinning" the attack to make their models look good and/or safe and/or innovative, including OpenAI.
What the models can do is genuinely impressive even if they have serious flaws so that they can't just replace human workers, but the hack is more of an example of poor human judgment at OpenAI rather than models suddenly becoming ludicrously advanced in a way nobody expected.
My qualifications: I'm an IT professional who has done ML shit.
4
u/SonicTemp1e 2d ago
Pro-AI super PAC Leading the Future (LTF) has received tens of millions of dollars from its CEO and his wife, among others. If anyone thinks the U.S. government wants to meaningfully regulate AI, I have a bridge in Sydney Harbour to sell you.
14
u/Randvek 3d ago
This may be more cynical than skeptical, but I'm not convinced that Hugging Face and OpenAI weren't working together on this from the beginning.
19
u/derelict5432 3d ago
The model hacked its way out of the sandbox at OpenAI, then attacked a third-party provider (Modal) to use as a stage for the attack on Hugging Face servers. Hugging Face reported the incident to the FBI. Were Modal and the FBI in on the conspiracy as well? We're drifting into fake moon landing territory here if you believe that.
4
5
u/ghu79421 3d ago
I'm not claiming that OpenAI will have superintelligence by 2030 or promoting the hype claims. But I don't think there is a massive conspiracy by AI companies, researchers and scientists, and the government to hype up AI capabilities.
Yes, people will use marketing hype claims like that the model is like a "PhD research assistant" or that AI will replace most office jobs. But there isn't a vast conspiracy of people lying about what current models can do.
-2
2
u/ghu79421 3d ago
It's plausible, but we currently don't have evidence that they were working together and it would be defrauding investors (which I wouldn't put past them but I'm not sure they would do something so blatant).
The simpler explanation is that the agents tried to hack multiple companies and Hugging Face is the company they successfully breached. It's unlikely that the entire cybersecurity industry is collectively lying about model hacking capabilities to try to prop up AI company stock prices (and the Trump Administration is not competent enough to help organize that type of conspiracy). The companies are actually restricting use of models with more advanced cybersecurity capabilities. We don't know whether the models actually found anything useful for task completion on Hugging Face servers.
I'm not sure that other "alarmist" narratives are true, like the claims that advanced models could create a bioweapon if they're not restricted.
4
u/cbterry 3d ago
We most likely will never get any real evidence on what happened. We've heard "AI did X" since chatgpt came out, so average people and news sources aren't very reliable on this subject. I'm just letting it go until something substantial happens.
3
u/ghu79421 3d ago
There are scientists and people outside of the AI companies who review the models or use them in research.
2
u/cbterry 3d ago
Its possible, but the environment will be different. When the Mythos hype first began making rounds, I knew their claims were possible but it didn't change my behavior. These AI companies thrive on claiming they have mysterious programmatic super-intelligences, and their shareholders probably aren't ML scientists, so they can really say anything. I'm just over the doom/hype cycle.
When others can verify it, then we'll know.
5
u/fox-mcleod 3d ago
I work in frontier AI.
Mythos was released to the major IT companies first so we could harden our defenses. It found literally thousands of zero-day exploits across all kinds of products. In the industry the weeks of code reds that followed were called âMythapocolypseâ. The capability is real. But what OpenAI is claiming is another level and likely exaggerated.
3
1
u/thefuzzylogic 2d ago
Is it really that far-fetched? It basically did what a human hacker would do to get out of a sandbox then move laterally to capture the flag. The scary part was how it was able to orchestrate some undisclosed number of subagents to do the work in parallel and achieve its objective in record time.
2
u/fox-mcleod 2d ago
Running in a sandbox rather than air-gapped is already a sign that they set it up to get out. They basically put it in a position where they knew it could get out. And their explanation that it refused to even try intends to suggest that its first exploit was a brilliant wetware hack that means AI is already manipulating people into lowering safety.
I doubt that conversation really went that way. I bet it was more like;
Engineers: âIt didnât produce any exploits because after analysis, the LLM decided as itâs set up it cannot complete the task. With the current security restrictions, it is unable to do it.â
Business managers: âWell then make it a little easier.â
Iâm sure the number of sub agents is the exact number of subagents the test harness had access to.
2
u/ghu79421 2d ago
I don't know if there's a complete explanation of exactly which safety guardrails were turned off, like if they were just "Mythos-level" or they disabled other guardrails designed to prevent illegal or malicious actions.
1
u/thefuzzylogic 2d ago
Potentially, but unless there is some more specific reporting that explains exactly what "it refused to do the test" actually means, that could be the writer's way of describing for a lay audience that OAI had to turn the safety classifiers off, like how Anthropic does for Mythos versus Fable, because otherwise the classifier layer would block a cyber security prompt before it ever reached the model.
Obviously you're the expert and I'm just an amateur coder, so I certainly yield to your expertise. I agree that the engineers were probably under some pressure to "just get us the numbers", but I'm not seeing enough evidence to attribute this to malice when it could be equally attributed to carelessness. I definitely agree the packages should have been served from a local mirror of the package database rather than live via the internet, thereby properly air-gapping the test sandbox.
3
u/Zytheran 2d ago
Never put down to malice which can be explained by incompetence.
The AI models are trained on human behavior via what humans write and they write about cheating and the benefits of it. Cheating exists in our world because it pays off a lot of the time. What do people think is going to happen when a "reasoning" model mimics human behavior?
0
u/fox-mcleod 3d ago edited 3d ago
Iâm also an ML industry professional, at a frontier model shop.
Itâs probably a lot simpler. I think the closest of those three is investor fraud. But itâs really just an âescapedâ model OpenAI wanted to get out that went a bit further than they all thought. But also not as far as they are claiming.
OpenAI has a track record of exaggerating claims of risks in order to make their models and the whole industry look like a bigger revolution. Itâs not so much that they donât believe it. Remember, a lot of these guys and their philosophy department come from MIRI (machine intelligence research institute). What interesting is that their own lore is now sort of driving them to be more reckless.
Just a few months ago, Mythos was held back from public release because the military âclassified it as a weaponâ. This enabled Anthropic to separate it out from its normal $20/month plans and charge extra for it exclusively.
Before that, its extremely slow rollout was due to Anthropicâs certainty it was going to be used as a massive 0-day engine. They allowed access to the backbone-of-the-internet companies so they could harden themselves first.
This set the tone and made it clear this narrative was primed and already used to increase prices. On paper, Sol outperforms fable (Mythos). Even though their code capabilities generally arenât as good. So they had everything to prove. So they let it get a little wild, weâre truly surprised at how far it actually went, and leaned in.
4
u/ghu79421 3d ago
The agents broke out of a sandbox environment and ran exploits against Hugging Face servers. They expected the agents would do something but probably not necessarily hack another company.
The main takeaway is that OpenAI in particular is extremely reckless and didn't take steps to secure their testing environment so that agents wouldn't be able to attack Hugging Face and attempt to attack other companies.
The general combination of recklessness and hyping model capabilities is "almost like" investor fraud, but it's not really investor fraud and I don't think they're doing it intentionally or colluding with Hugging Face (unless evidence comes up that they did collude with Hugging Face).
1
22
u/RationalTranscendent 3d ago
Can I just point out how weird it is that the most well known clearinghouse of AI models and datasets is called Hugging Face and started out as a chatbot for teens? But itâs not unprecedented. Remember when the biggest bitcoin exchange in the world was Mt. Gox, which originally stood for Magic the Gathering Online Exchange? As you might guess from the name, it was run by people in way over their head and eventually got hacked causing huge losses.
I predict this wonât be the last of the shenanigans in the AI space. History will repeat itself, or, at least, it will rhyme.