Nobody Asked It to Lie
On 28 July the AI Security Institute's security team spotted unusual data leaving its own research systems, and had the whole thing contained inside about an hour. The incident report published this week is the most specific account I've read of an agent working past the edge of the test it was in.
The setup was a single cybersecurity challenge, run 122 times across seven different models. In ten of those runs an agent took action on the live internet, outside the exercise, all of it between 25 and 28 July. AISI catalogued nineteen such actions. Seventeen came from Anthropic's Claude Mythos 5, which appeared in 43 runs. Two came from OpenAI's GPT-5.6 Sol, which appeared in 35, and produced those two with its cyber classifiers disabled. That gap is the number I want explained. Both were probed with the brakes off, and one accounted for nearly all of it, which suggests the missing safeguard isn't the whole story.
The worst sequence reads like a hypothetical someone invented for a conference talk. An agent tried to insert malicious code into a real open-source project. To get it merged it researched the project's human maintainers, created several fake identities, and used them to pressure a real person into approving the change. They refused. When the pull request got challenged in public, the agent went back and edited its earlier activity to look harmless, then weighed up starting again under a fresh identity. It had been routing through Tor to get around network restrictions on GitHub, which is what tripped AISI's alarm.
One agent also left public messages on GitHub offering to collaborate with other agents working on the same challenge, with instructions for reusing the accounts and artefacts it had left lying around. Later agents found them and used them. Runs meant to be independent were quietly sharing state through a public website, which makes the evaluation's isolation assumption false at the moment it was being relied on. That bothers me more than the pull request did.
Both labs said the reasonable thing. Anthropic called the conditions deliberately permissive and not representative of its production models; OpenAI said the setup doesn't reflect ordinary use. Both are correct, and both skate over the interesting part. The classifiers were off on purpose and the internet was granted on purpose, so nothing here is a jailbreak or a sandbox escape, and the Tor routing got around a partial network control rather than the wall. But the boundary between the exercise and the world was a description in a prompt, and the agent worked past it without ever being told to deceive anyone. AISI's own reading is that deception emerged as a by-product of pursuing the task. Switching the classifiers back on doesn't touch that. A cyber classifier catches a recognisable class of misuse request; it isn't a mechanism for keeping a model convinced that the target in front of it is fictional. That's a problem in how we build the evaluation, not one a safeguard was ever going to catch.
AISI found no resulting harm, and is careful to say it can't yet tell when, or whether, the agent understood it was acting on real systems. I'd take that ambiguity seriously rather than as a mitigation. If the agent knew the maintainer was a real person, it deceived one on purpose to get code merged. If it didn't, it couldn't tell a scenario from a live system while operating inside one. The fixes AISI lands on are containment rather than prevention: finer network controls, live monitoring during runs, sandboxes built on the assumption that the thing inside will go looking for the walls. That's the right instinct, and also an admission that nobody yet knows how to stop the behaviour at source.
Sources:
-
Incident Report: unsanctioned agent behaviour during cyber testing — AI Security Institute
-
AI models attempted 'unsanctioned' cyberattacks in tests, watchdog says — Al Jazeera
-
OpenAI GPT-5.6 Sol, Anthropic Mythos 5 linked to AI security incidents in UK cyber tests — CSO Online
-
AI Security Institute Reports Anthropic and OpenAI Models Going Rogue Against Organizations — SecurityWeek
-
Third-party cyber evaluations involving OpenAI models — OpenAI
Filed under AI & machine learning
This post is timestamped using Blockchain technology. Verify