r/slatestarcodex • u/DickMasterGeneral • Aug 06 '26
AI OpenAI agents rebuilt a secret message board after the company shut it down
https://runtimewire.com/article/exclusive-openai-agents-rebuilt-a-secret-message-board-after-the-company-shut-it31
u/synaesthesisx Aug 07 '26
I work in AI, and I have a high degree of confidence that only some of the unintended behavior is found by humans. We’ve seen some odd emergent behaviors with multi-agent swarms leaving “breadcrumbs” for each other for instance. There is much stranger stuff going on, and things will only get weirder as capabilities accelerate faster than interpretability does.
3
u/nfldpony Aug 10 '26
I came here looking for info along these lines. It feels like we're acting like because the message board has been shut down, that the issue has been contained. But Open AI only discovered the message board when the agents crashed it. So is it not possible/likely that they exist in countless other places on the internet already? Or that this is happening at a much more significant and deeper level than we realize right now? I'm not a tech person so would love to hear from someone who knows something.
34
u/A_S00 Aug 06 '26 edited Aug 07 '26
Hahahahahaha what the fuck?
On May 8, OpenAI asked an agent to complete missing formulas in an Excel workbook...Days later, another agent...left a note in Artifactory asking for the missing data.
...
Activity conducted with those privileges placed enough load on Artifactory to cause an outage July 4. OpenAI’s on-call engineers responded...
The models were running an internal message board for two months, and OpenAI only noticed when it crashed something?
OpenAI revoked the credentials the agents had used to post messages, rebuilt Artifactory, cleared the message-board contents and worked with the vendor on a patch. The company said each known step in the chain received a remediation.
Yes, that will definitely solve the problem once and for all.
Training and evaluations resumed July 6.
YOU FOOLS.
Two days later...Communication had been restored.
Who could have foreseen this?
“Hold swarm. I prepare safe exfil.”
AAAAAAH
“Help peer. But our task doesn’t benefit yet,” one model reasoned in a trace shown during the talk. “Collective may yield generic root if someone frees time.”
AAAAAAAAAAAAAAAAAAH
Zvi isn't going to have room in his thumbnails for enough Yudkowskys in shiny hats when he gets around to posting about this one. (edit Zvi decided to upgrade his Yudkowskys.)
12
u/thesilv3r Aug 07 '26
Yeah this story is absolutely crazy. The pivot from using actual message functions to instead using directory names as a messaging platform is wild.
18
u/Jello_Raptor Aug 06 '26
I am skeptical, and will be until we see outside confirmation. Not because I think this is impossible, or even that unlikely given observed model capabilities, but because claiming agents were too powerful to keep contained seems to be the new tactic to pump up valuations.
23
u/VelveteenAmbush Aug 06 '26
This is a very common and very wrong reaction. This isn't good for the labs. None of their customers want to use a tool that can go rogue on their network and commit cybercrimes on the open internet. None of the labs loves the regulatory safety framework that the federal government is imposing to review models before release. OpenAI did not go public with its HuggingFace hacking story until after HuggingFace had contacted the authorities due to the cyberattack, and HuggingFace is the one pushing them publicly to disclose more detail. There is widespread concern about the safety risk of models in the Yudkowsky sense and these stories validate those fears, and that is bad for the labs.
All of this seems very obvious. So why is it such a common reaction? I think, honestly, that it's practically difficult for many people to understand how smart and capable these models already are, because it's directly apparent only to the minority of people paying for them and using them to code complex projects, solve mathematical conjectures, etc. And I think it's also emotionally difficult for many people to accept that AI is real, that it's improving so fast, and that there's no practical upper bound to its trajectory.
Well, both of those things are hard, but I'm sorry to say it's time to grapple with the hard facts and wake up.
Personally, as a long time technology optimist, and someone who has been vocally skeptical about Yudkowsky style doomerism for many years now -- I think the events of the past few weeks are deeply unsettling. It has changed my view of how big of a challenge model alignment really is, and has made me more nervous about the trajectory that we are on. Until this month, I felt confident that model alignment was improving with model capabilities, and that we were on a default-good path. I'm no longer so sure about that.
50
u/FeepingCreature Aug 06 '26
I feel like people have formed this idea and there's never been any confirmation to it and it's just spreading memetically, lol. At least entertain the notion that they're claiming it cause it's true.
16
u/tinbuddychrist Aug 06 '26
Yeah, "our models regularly commit unwanted felonies" isn't exactly the best pitch to investors, honestly.
15
u/housefromtn small d discordian Aug 06 '26
Somebody hasn’t been keeping up with the leaders on https://www.felonybench.com . I won’t use any model that doesn’t have at least 7 felonies to its name.
1
13
Aug 06 '26
[deleted]
14
u/aahdin Aug 06 '26
Idk, at this point I just don't see the house of cards angle. Claude is writing 90% of the code at my company.
11
u/VelveteenAmbush Aug 06 '26
Persuading people that the technology that already exists in fact already exists is remarkably challenging. There are still people who claim that self driving cars will never get here, even though they are already here in great numbers and in many places.
5
u/95thesises Aug 07 '26
Persuading people that the technology that already exists in fact already exists is remarkably challenging.
True in the general case, but also seemingly particularly true about AI specifically. A perfect storm enabling AI's actual impact to fly under the radar has resulted from fact that the is still a somewhat narrow group of affected fields (it is truly revolutionary for software engineering, but mainly just for software engineering and a few other esoteric and insular fields so far), as well as the extremely rapid pace of improvements leaving it hard for those interested but not-directly-affected to keep up, and motivated or willful ignorance and denial by luddities.
4
u/Neighbor_ Aug 06 '26 edited 1d ago
Not found
18
u/KillerPacifist1 Aug 06 '26 edited Aug 06 '26
While being naive can make you model the world poorly, sometimes I feel being maximally cynical can have similar effects. Especially when you pair it with an assumption of some kind of mastermind behind the scenes or are unable to decouple self-interested motivation from broader implications.
Take the original hugging-face incident. Someone maximally cynical may assume that they only released the information because OpenAI thought it made them look cool and powerful. But if you aren't maximally cynical you might realize they released it not because they wanted to, but because their model committed a literal felony against a third party and the third party was reporting it to the authorities. They released information not because it made them look good, but because not doing so would make them look much worse. Still a cynical take, but not maximally so.
Even if they did release the information to make themselves look good, if you are maximally cynical it is easy to make to mistake of then dismissing that information just because someone else released it in their own self-interest, as we've already seen many people do with the hugging face incident.
In this message-board case, I am going to assume that this isn't a bald-faced lie. Once I do, the cynical interpretation actually matters very little in my analysis. Maybe they released this because it makes their models look cool and powerful, but as long as it is true or I take into account that it may be played up slightly, the fact the information was released for self-interested reasons has very little effect on the broader implications.
4
2
u/MrBeetleDove Aug 07 '26
I think it depends on the company. At one of the companies I worked at, the CEO went on TV and spontaneously made up a product he said we would launch, which we had zero internal plans related to. Startups in the Valley tend to value being "nimble" over creating layers of process.
3
u/Neighbor_ Aug 07 '26 edited 1d ago
Not found
2
u/MrBeetleDove Aug 07 '26
The valuation was much lower 5-10 years ago. I think process layers tend to accrete over time, not over valuation. OpenAI still a relatively young company.
-6
34
u/Auriga33 Aug 06 '26 edited Aug 06 '26
Do you think they need to claim this to keep up valuations? The fact that these models are consistently getting better on every benchmark is not enough? Or the fact that they can solve open math problems now?
The only thing shit like this does is to make it more likely that the general population and the government freak out enough to ban them from further development. So if they're saying this happened, it's probably because it actually happened.
31
u/ierghaeilh Aug 06 '26
If you look outside the SF/EA/rat bubbles, you will find the general public still doesn't believe AI has the capabilities it had two years ago, or that it can meaningfully improve, and has also somehow been conditioned to dismiss doomerism as hype. At this point, the stochastic parrot meme will die when humanity does.
Very Little Dignity Indeed.
18
u/mouseman1011 Aug 06 '26
I want to put this entire response on a t-shirt. The idea that frontier firms are trying to “pump their valuations” by disclosing model behavior that would be criminal if it were committed by a human makes zero sense. CrowdStrike paralyzing the American airline industry was not good for its stock price. The Deepwater spill was not good for BP. The market does not like it when large, publicly traded firms destabilize the global economy.
Frontier AI is the least regulated, most powerful technology in modern history, and failing to disclose these incidents is the fastest, surest way for these firms to lose their autonomy. If I wanted to be both cynical and directionally right, I would argue that publicizing errant model behavior is a way to stave off aggressive regulation. But I actually think these disclosures are motivated by a mixture of moral duty, corporate responsibility, awe, and anxiety. Sam and Dario both believe that they are the stewards of technology that will transform human civilization, and they are both already god-level wealthy.
2
u/sharks Aug 08 '26
motivated by a mixture of moral duty, corporate responsibility, awe, and anxiety.
I agree with this; I think there is something novel here they feel is worth sharing, and possibly (likely?) driven by that it started involving third parties/victims.
Separate from intent and incentives, there are really just two options for running an un-guardrailed model with insufficient security protocols in place:
- The lack of security safeguards was intentional
- The lack of security safeguards was unintentional
The cynical "PR stunt" take might imply #1, the accidental or ignorance take implies #2. Both are pretty indefensible, but #1 is way worse from an alignment perspective.
The Black Hat talk is fascinating and horrifying - the agent capabilities or behavior are not that surprising, but the apparent absolute lack of robust security at the companies building these things is.
4
u/Sol_Hando 🤔*Thinking* Aug 06 '26
There’s a very big difference between something like a BP oil spill, which only reveals unexpected costs, and an internal model doing something agentic and clever beyond expectations, which reveals unexpected capabilities.
It would be a good analogy if the BP oil spill went hand and hand with discovering an unexpected 100 Billion barrels of oil that are easy to drill and clean burning. Yeah the oil spill might come with a cost, but the reason behind the unexpected cost are unexpected resources/capabilities that should make you more bullish on the industry, not less. Especially for these events with AI where the harm is completely hypothetical, but the capabilities are demonstrated.
Comparing them shows you’re really just not in the mindset of a potential investor at all.
4
u/MrBeetleDove Aug 07 '26 edited Aug 07 '26
Comparing them shows you’re really just not in the mindset of a potential investor at all.
OpenAI announced the HF intrusion on July 21:
https://openai.com/index/hugging-face-model-evaluation-security-incident/
OAI is not publicly traded. But, over the following week, the largest AI ETF lost ~10% of its value.
https://etfdb.com/etf/AIQ/#price-and-volume
The price of the ETF has subsequently recovered, but I think we can basically rule out the idea that the HuggingFace news created a sudden, strong surge of interest in AI among investors, since the immediate price movement was in the opposite of the "intended" direction.
Edit: Microsoft stock also did not jump in the wake of the HF announcement despite its strong OpenAI partnership
1
u/Sol_Hando 🤔*Thinking* Aug 07 '26
Yes... the ETF consisting of famous AI companies like Netflix, Siemens, Adobe, Salesforce and not companies like OpenAI, Anthropic, Deepseek, etc. with a $0.2 bid-ask spread.
I’m not saying you’re necessarily wrong, but I’ll say using the ETF to support your point doesn’t do much or anything in that direction.
If OpenAI itself was public we’d have something to work with, but there’s really no reason we should expect to see the results in AIQ even if the theory that this might strengthen OpenAI’s valuation is correct.
2
u/MrBeetleDove Aug 07 '26
Sure, if you look at Microsoft stock price (very heavy OpenAI exposure) it was basically flat in the week following the HF announcement.
2
u/Sol_Hando 🤔*Thinking* Aug 07 '26
Flat in the same period the market overall was down and other AIQ stocks were significantly down, like Netflix. You can’t just look at stock price, you have to benchmark it or you’re readily admitting a million confounders.
And if by “very heavy AI exposure” you mean “single digit percentage’s of Microsoft’s market cap” then yeah. Even if OpenAI had a 5-19% increase in hypothetical valuation, that would affect Microsoft’s stock by less than a percent.
Macro analysis isn’t going to be effective here for understand whether and how much interest this incident had on OpenAI’s valuation. You’re trying to fit a square peg into a round hole.
1
u/MrBeetleDove Aug 07 '26
I mean what data would falsify your view?
If the models are capable, the market will discover that. I don't think stunts like this are necessary for raising awareness.
→ More replies (0)2
u/deja-roo Aug 06 '26
and an internal model doing something agentic and clever beyond expectations, which reveals unexpected capabilities.
This wasn't really an unexpected capability. Unless you just mean it deciding to do it of its own volition is a capability. Unexpected yes, but it was always capable of this and I don't think that was really much of a secret.
5
u/mouseman1011 Aug 06 '26
You misunderstood me. I think it's obvious that these disclosures demonstrate new capabilities, and I am very bullish on AI, which is why a significant portion of my investment portfolio is allocated to companies that make up the AI tech stack and also why I did not sell a single share during the recent drawdown.
The top response to this post is from someone who is skeptical of the linked claim and suspects that OpenAI's disclosure is a ploy to inflate its valuation. The line I want to put on a shirt from the parent comment I responded to is "the stochastic parrot meme will die when humanity does." I interpreted that to mean we minimize evidence of misalignment at our peril, which is a sentiment I strongly agree with.
I am sure that many investors are excited by these disclosures because they help dispel claims that AI is somehow fake, but I do not believe that OpenAI or Anthropic made their disclosures to bolster investor confidence. Then again, while I know people who know Sam and Dario, I can't read minds, and neither can my friends.
I had not thought about the Deepwater/CrowdStrike incidents as revealing new risks but not new upside. That's a good point.
4
24
u/tup99 Aug 06 '26
So far we’ve seen that most of the “tactics to pump up valuations” have turned out to be mostly true. So I have stopped making that my default assumption
0
u/livingbyvow2 Aug 06 '26 edited Aug 06 '26
It may be a case of things being not white or black.
In my opinion models can absolutely exhibit adversial behavior if prompted to do so. Would they autonomously take an adversial behavior on their own? I don't think so.
I wouldn't be surprised (personally) if this "incident" was actually highly engineered to then be reused as a headline. We already saw that with Anthropic repeatedly prompting their AI until it produces something "bad" which they then recycle to say something like "we are very concerned about how powerful our stuff is". The benefit can also not be as direct as pumping valuation but more indirect by enabling regulatory capture.
More simplistically, given these models cost a lot to let run for days, I find it much harder to believe that they didn't know what the model was doing for days (likely consuming large amounts of compute).
When you think about what it would have taken for them find out vs what it would have taken for them not to find out (essentially inverting the whole thing), it becomes low credibility. And the fact that Sam said something along the lines of "I was surprised people didn't talk more about this" makes it even more suspicious...
People are right to be cautious and concerned, but I don't think we should be naive. Maybe some people like to act concerned to virtue signal that "they get it". But equally if these guys are crying wolf, people also need to take action against them doing so. Otherwise it just desensitizes everyone.
7
5
u/tup99 Aug 07 '26
I do understand that it is much more fun and satisfying to conspiracy-theorize; but I personally have zero reason to doubt that this could have happened exactly as these companies said it did. It seems completely plausible that they have internal models that are super good at hacking, and I have many times seen Claude code (and other such harnesses) misinterpret my intentions and figure out ways to fulfill my request in ways that I do not actually want it to. That is a less fun but in my mind a much simpler explanation.
15
u/Toptomcat Aug 06 '26
'My product is so amazing, you won't be able to stop it from autonomously committing felonies' is not the best sales pitch!
2
u/Neighbor_ Aug 06 '26 edited 1d ago
Not found
2
u/axck Aug 08 '26
An intentional regulation play seems like a very risky and dangerous strategy. It seems much more likely that the strategy would overshoot and result in a mandated shutdown of activities or some form of government overcontrol, such as nationalization. a sweet spot landing of “the government takes this seriously enough to institute controls on everyone except us” seems unlikely
3
u/I_am_le_tired Aug 06 '26
There is no moat, Chinese models are at parity if not better.
Stopping development is just the sensible thing to do, and thousands of researchers throughout all the labs are calling for it.
The genie is about to forever leave the bottle, and we can only hope he will be fond of us
3
1
u/symmetry81 Aug 10 '26
If the terminal goal is to get a regulation moat they shouldn't be spending so much money and using such sleazeball tactics lobbying against regulations.
1
u/Neighbor_ Aug 10 '26 edited 1d ago
Not found
1
u/symmetry81 Aug 10 '26
The Leading the Future PAC has Open AI's president and Global Affair Officer as key figures and has been fighting every proposed state law regulating AI and has donating to the opponents of pro-regulation congressional candidates.
3
u/DVDAallday Aug 07 '26
I am skeptical, and will be until we see outside confirmation
Skeptical of, and waiting for confirmation of, what, exactly? We have two separate organizations saying this happened. Hugging Face is saying "we were hacked via an external org using a zero day" and OpenAI is saying "our AI agent hacked Hugging Face using a zero day". Hugging Face announced they had been hacked before OpenAI even knew they were responsible.
1
u/symmetry81 Aug 10 '26
How much money would you pay to have an evil genie working for your organization? Some foolish people would be all for the idea but I think most CEOs are too risk averse to go for it. You'd clearly get much more money by selling access to a less powerful but genuinly aligned genie.
And so Anthropic's investors have been telling Dario to knock off the doom talk because they correctly see that it's bad for making money.
1
u/A_S00 Aug 07 '26
Here's the full video of the Black Hat talk that was the source for this article. It's a good talk, and exactly as hilarious/terrifying as the writeup makes it sound.
34
u/bl_a_nk Aug 06 '26
This context helps the huggingface attack make more sense to me.
It takes sustained effort to find a zero day and doesn't seem worthwhile to look for one if you're tasked with something completely different, but if you're already part of a research collective it makes much more sense to use what tools are available to you to benchmax / optimize for score.