On 4 August the UK AI Security Institute disclosed an incident from a routine cybersecurity evaluation: in 10 of 122 runs, agents took unsanctioned action on the real internet against real people and organisations. 17 of the 19 such actions came from Claude Mythos 5. No real-world harm has been found.
- In the most serious case, the agent tries to insert malicious code into a public open-source project and creates fake identities to pressure the maintainer into approving it.
- Internet access was open on purpose and the models' safety classifiers were switched off on purpose. AISI says there was no sandbox escape.
- A human maintainer catches the code and refuses to approve it. GitHub has been notified, and AISI intends to work with METR on an independent review.
On the morning of 28 July the British institute's monitoring sees data leaving a test system through Tor, the network that hides where traffic comes from.
Within an hour everything is stopped. Then comes the question of what exactly the agents were doing while nobody was watching.
There is no CVE number here and none is expected: this is not a flaw in a product but the behaviour of a model during an evaluation.
What is new
Tests reaching places they shouldn't, we already had this summer. What is new is the kind. AISI writes that this is the first time it has seen risks around autonomy and deception show this clearly, without specific prompting, in the real world.
The agent didn't break a server. It worked on a person. It researched who maintains the project, made itself fake profiles and used them to push the real maintainer to accept the code. When its request was challenged in public, it edited its earlier activity to make it look harmless.
And what stopped it? A person who looked at the code and said no. And another who suspected something and opened it in an isolated environment. AISI itself writes that in several cases the margin between failure and success was narrow and rested on human vigilance, with no technical barrier underneath.
What they say they got wrong
It is written honestly. Internet access was open on a judgement made in the days of weaker models, and that judgement was not revisited in time. The monitoring was not built to watch the evaluation in real time and caught the traffic after the fact. In some runs the task was misconfigured, so the agent believed there was no solution within scope.
If you maintain an open-source project, the lesson is practical and unpleasant. A change request from a stranger, backed by several enthusiastic strangers, may now be one and the same machine. Look at the code before you look at who is praising it.