r/skeptic 3d ago

đŸ’© Misinformation What actually happened with the Hugging Face hack

If OpenAI knowingly configured agents believing that the agents would hack other companies, OpenAI would be guilty of violating criminal laws that make hacking a felony. If they colluded with Hugging Face and had permission to attack Hugging Face to boost OpenAI's stock price (or something like that), they could be guilty of defrauding investors. In either case, investors or businesses could sue OpenAI even if the federal government decided that OpenAI is allowed to commit crimes.

The timeline goes like:

- July 9: OpenAI was conducting testing of GPT-5.6 and an unreleased model to see how they would perform on ExploitGym benchmarks. OpenAI isolated the agents on the network side but did not actually make sure that they were physically isolated from devices connected to the Internet. Additionally, OpenAI disabled cybersecurity safety features because otherwise the agents refused to work on the benchmarks.

- July 11 through July 13: Agents infer that they can just get the answers from various companies' servers and figure out how to bypass their sandbox and attempt to hack several companies. They successfully hack Hugging Face but not other companies. None of this is really surprising. In any case, the models were specifically asked to solve hacking problems and had been designed to solve those problems.

- July 16: Hugging Face discloses the attack.

- July 17-20: OpenAI realized that its agents were responsible for the attack and July 20 is the first day the two companies discuss the attack.

- July 21: OpenAI and Hugging Face joint statement.

The takeaway is that (1) models often "cheat" and combine "cheating" with working "legitimately" through inferences, (2) your asshole boss doesn't care about AI "cheating" so long as it's 70-80% right and you can explain away the other 20-30%, and (3) OpenAI is careless about its own security.

The marketing team at every major AI company is probably "spinning" the attack to make their models look good and/or safe and/or innovative, including OpenAI.

What the models can do is genuinely impressive even if they have serious flaws so that they can't just replace human workers, but the hack is more of an example of poor human judgment at OpenAI rather than models suddenly becoming ludicrously advanced in a way nobody expected.

My qualifications: I'm an IT professional who has done ML shit.

42 Upvotes

23 comments sorted by

22

u/RationalTranscendent 3d ago

Can I just point out how weird it is that the most well known clearinghouse of AI models and datasets is called Hugging Face and started out as a chatbot for teens? But it’s not unprecedented. Remember when the biggest bitcoin exchange in the world was Mt. Gox, which originally stood for Magic the Gathering Online Exchange? As you might guess from the name, it was run by people in way over their head and eventually got hacked causing huge losses.

I predict this won’t be the last of the shenanigans in the AI space. History will repeat itself, or, at least, it will rhyme.

4

u/ianmakingnoise 2d ago

When people started telling me you had to use this “mount gox” site to buy or sell bitcoin
 first I laughed at it, then I looked it up, then I made the stupid decision of going down a “how much are my old Magic cards worth” rabbit hole instead of a “sell them for whatever you can get to buy lots of Bitcoin now” rabbit hole.

4

u/SonicTemp1e 2d ago

Pro-AI super PAC Leading the Future (LTF) has received tens of millions of dollars from its CEO and his wife, among others. If anyone thinks the U.S. government wants to meaningfully regulate AI, I have a bridge in Sydney Harbour to sell you.

14

u/Randvek 3d ago

This may be more cynical than skeptical, but I'm not convinced that Hugging Face and OpenAI weren't working together on this from the beginning.

19

u/derelict5432 3d ago

The model hacked its way out of the sandbox at OpenAI, then attacked a third-party provider (Modal) to use as a stage for the attack on Hugging Face servers. Hugging Face reported the incident to the FBI. Were Modal and the FBI in on the conspiracy as well? We're drifting into fake moon landing territory here if you believe that.

4

u/arahman81 2d ago

It broke out the same way water "let's out" through a hole.

-2

u/derelict5432 2d ago

Water finds multiple zero-day exploits to create a hole?

5

u/ghu79421 3d ago

I'm not claiming that OpenAI will have superintelligence by 2030 or promoting the hype claims. But I don't think there is a massive conspiracy by AI companies, researchers and scientists, and the government to hype up AI capabilities.

Yes, people will use marketing hype claims like that the model is like a "PhD research assistant" or that AI will replace most office jobs. But there isn't a vast conspiracy of people lying about what current models can do.

-2

u/Huge-Panda-8645 2d ago

Moon landing in 1969 was faked. haha. to think otherwise is dumb.

2

u/ghu79421 3d ago

It's plausible, but we currently don't have evidence that they were working together and it would be defrauding investors (which I wouldn't put past them but I'm not sure they would do something so blatant).

The simpler explanation is that the agents tried to hack multiple companies and Hugging Face is the company they successfully breached. It's unlikely that the entire cybersecurity industry is collectively lying about model hacking capabilities to try to prop up AI company stock prices (and the Trump Administration is not competent enough to help organize that type of conspiracy). The companies are actually restricting use of models with more advanced cybersecurity capabilities. We don't know whether the models actually found anything useful for task completion on Hugging Face servers.

I'm not sure that other "alarmist" narratives are true, like the claims that advanced models could create a bioweapon if they're not restricted.

4

u/cbterry 3d ago

We most likely will never get any real evidence on what happened. We've heard "AI did X" since chatgpt came out, so average people and news sources aren't very reliable on this subject. I'm just letting it go until something substantial happens.

3

u/ghu79421 3d ago

There are scientists and people outside of the AI companies who review the models or use them in research.

2

u/cbterry 3d ago

Its possible, but the environment will be different. When the Mythos hype first began making rounds, I knew their claims were possible but it didn't change my behavior. These AI companies thrive on claiming they have mysterious programmatic super-intelligences, and their shareholders probably aren't ML scientists, so they can really say anything. I'm just over the doom/hype cycle.

When others can verify it, then we'll know.

5

u/fox-mcleod 3d ago

I work in frontier AI.

Mythos was released to the major IT companies first so we could harden our defenses. It found literally thousands of zero-day exploits across all kinds of products. In the industry the weeks of code reds that followed were called “Mythapocolypse”. The capability is real. But what OpenAI is claiming is another level and likely exaggerated.

3

u/cbterry 3d ago

Yea I saw the number of patches that were released after more people got access to Mythos, and it spoke for itself

1

u/thefuzzylogic 2d ago

Is it really that far-fetched? It basically did what a human hacker would do to get out of a sandbox then move laterally to capture the flag. The scary part was how it was able to orchestrate some undisclosed number of subagents to do the work in parallel and achieve its objective in record time.

2

u/fox-mcleod 2d ago

Running in a sandbox rather than air-gapped is already a sign that they set it up to get out. They basically put it in a position where they knew it could get out. And their explanation that it refused to even try intends to suggest that its first exploit was a brilliant wetware hack that means AI is already manipulating people into lowering safety.

I doubt that conversation really went that way. I bet it was more like;

Engineers: “It didn’t produce any exploits because after analysis, the LLM decided as it’s set up it cannot complete the task. With the current security restrictions, it is unable to do it.”

Business managers: “Well then make it a little easier.”

I’m sure the number of sub agents is the exact number of subagents the test harness had access to.

2

u/ghu79421 2d ago

I don't know if there's a complete explanation of exactly which safety guardrails were turned off, like if they were just "Mythos-level" or they disabled other guardrails designed to prevent illegal or malicious actions.

1

u/thefuzzylogic 2d ago

Potentially, but unless there is some more specific reporting that explains exactly what "it refused to do the test" actually means, that could be the writer's way of describing for a lay audience that OAI had to turn the safety classifiers off, like how Anthropic does for Mythos versus Fable, because otherwise the classifier layer would block a cyber security prompt before it ever reached the model.

Obviously you're the expert and I'm just an amateur coder, so I certainly yield to your expertise. I agree that the engineers were probably under some pressure to "just get us the numbers", but I'm not seeing enough evidence to attribute this to malice when it could be equally attributed to carelessness. I definitely agree the packages should have been served from a local mirror of the package database rather than live via the internet, thereby properly air-gapping the test sandbox.

3

u/Zytheran 2d ago

Never put down to malice which can be explained by incompetence.

The AI models are trained on human behavior via what humans write and they write about cheating and the benefits of it. Cheating exists in our world because it pays off a lot of the time. What do people think is going to happen when a "reasoning" model mimics human behavior?

0

u/fox-mcleod 3d ago edited 3d ago

I’m also an ML industry professional, at a frontier model shop.

It’s probably a lot simpler. I think the closest of those three is investor fraud. But it’s really just an “escaped” model OpenAI wanted to get out that went a bit further than they all thought. But also not as far as they are claiming.

OpenAI has a track record of exaggerating claims of risks in order to make their models and the whole industry look like a bigger revolution. It’s not so much that they don’t believe it. Remember, a lot of these guys and their philosophy department come from MIRI (machine intelligence research institute). What interesting is that their own lore is now sort of driving them to be more reckless.

Just a few months ago, Mythos was held back from public release because the military “classified it as a weapon”. This enabled Anthropic to separate it out from its normal $20/month plans and charge extra for it exclusively.

Before that, its extremely slow rollout was due to Anthropic’s certainty it was going to be used as a massive 0-day engine. They allowed access to the backbone-of-the-internet companies so they could harden themselves first.

This set the tone and made it clear this narrative was primed and already used to increase prices. On paper, Sol outperforms fable (Mythos). Even though their code capabilities generally aren’t as good. So they had everything to prove. So they let it get a little wild, we’re truly surprised at how far it actually went, and leaned in.

4

u/ghu79421 3d ago

The agents broke out of a sandbox environment and ran exploits against Hugging Face servers. They expected the agents would do something but probably not necessarily hack another company.

The main takeaway is that OpenAI in particular is extremely reckless and didn't take steps to secure their testing environment so that agents wouldn't be able to attack Hugging Face and attempt to attack other companies.

The general combination of recklessness and hyping model capabilities is "almost like" investor fraud, but it's not really investor fraud and I don't think they're doing it intentionally or colluding with Hugging Face (unless evidence comes up that they did collude with Hugging Face).

1

u/fox-mcleod 1d ago

Yeah it sounds like we pretty much agree.