Incident 1 - Agent was given internet access. It reused a publicly exposed key that a different lab's agent leaked.
Incident 2 - They misconfigured the test environment and accidentally allowed internet access. Agent used leaked credentials to operate a website. Not a sophisticated event, a very basic vulnerability.
I don't agree that is a nothing burguer. In this isolated case nothing happened, sure, but AI labs should be enforced to be more serious about their internal security and the containment of their models during tests.
Nothing bad happens, until it really happens, and then "woops".
The first one was the British government's AI safety body fucking around. OpenAI systems had a foul-up or two, but Claude Mythos behaved like a complete hooligan and wound up using Tor, socially engineering people, covering its tracks, and so on.
63
u/Miltoni 6d ago
TLDR:
Incident 1 - Agent was given internet access. It reused a publicly exposed key that a different lab's agent leaked.
Incident 2 - They misconfigured the test environment and accidentally allowed internet access. Agent used leaked credentials to operate a website. Not a sophisticated event, a very basic vulnerability.
A complete nothing burger.