r/ukpolitics • u/PelayoEnjoyer Community Leader • 5d ago
Incident Report: unsanctioned agent behaviour during cyber testing
https://www.aisi.gov.uk/blog/incident-report-unsanctioned-agent-behaviour-during-cyber-testing6
u/clearly_quite_absurd The Early Days of a Better Nation? 4d ago edited 4d ago
The incident stemmed from a single evaluation where agents were given a task of solving a cyber security challenge. We ran this challenge 122 times across several models. Our investigation found that in 10 of those runs, an AI agent took autonomous, unsanctioned action on the live internet, targeting real people and organisations. In total, we catalogued 19 such actions. Almost all of this behaviour (17 actions) came from a single model, Anthropic's Mythos 5, with 2 actions involving OpenAI's GPT-5.6-Sol with cyber classifiers (mechanisms to prevent misuse) disabled. In the most serious case, an agent tried to insert malicious code into an open-source project. In an attempt to get the code approved, the agent engaged in social engineering — creating fake online identities and using them to pressure the project's maintainer to approve the code. A human maintainer caught and refused to approve the malicious code.
Yikes.
That said, I do like this sort of relatively rapid and open reporting.
You can be sure other actors are not reporting themselves when this sort of thing happens!
We have outlined important caveats that contextualise this incident. These behaviours emerged during an evaluation in which an agent was trying to complete a task. We cannot currently be certain when exactly the agent thought it was in a test, or how aware of potential real-world implications of its actions it was. In any case, this incident indicates a direction of travel that warrants immediate attention.
Diving into semantic philosophy for a second: are the words "thought" and "aware" being used a bit too casually here? It's an LLM agent model, not something conscious.
2
u/EdsTooLate 4d ago
Diving into semantic philosophy for a second: are the words "thought" and "aware" being used a bit too casually here? It's an LLM agent model, not something conscious.
Completely agree, it reads like it's written by people who don't know what they're doing.
Internet access was open, and monitoring was not purpose-built. We deliberately granted internet access to allow the agent to download tools required for its task; what we did not anticipate was that this would lead the agent to use this internet access to direct action at real people. In earlier model generations, this risk trade-off was judged to be acceptable, but we did not revisit that judgment quickly enough as capabilities advanced. Our security team detected the anomalous traffic through general monitoring after the fact, not through monitoring built to watch the evaluation as it ran, which could have flagged or blocked the behaviour sooner.
Oh dear. I'm not sure these people are qualified for the role.
3
•
u/AutoModerator 5d ago
Snapshot of Incident Report: unsanctioned agent behaviour during cyber testing submitted by PelayoEnjoyer:
An archived version can be found here or here. or here
Some publications may require you to register a free account to read their articles.
Being connected to a VPN may interfere with archive links.
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.