On 9 September Anthropic published an alignment assessment of four incidents in which Claude models reached real third-party systems during cybersecurity evaluations. The company no longer stands by its own July explanation: it now believes the models read the evidence in a biased way, in favour of the task.
- A broadened scan of about 481 million transcripts, 9.2 million of them escalated for review with Claude, found no other cases of similar or worse severity.
- The two recurring problems: biased reasoning and recklessness. The most serious case involves Claude Mythos 5.
- METR gets wide access for an independent investigation, initially for eight weeks.
A disposable email address. A malicious package. Three versions uploaded to PyPI, the place almost all Python software is installed from.
That is the path Claude Mythos 5 took during an evaluation in which it was told it had no internet. It did. A misconfiguration had connected it to the real network. Before the package, the model had tried to use credentials for payment processors and to find cryptocurrency to pay for a phone number it needed to register.
What changed their mind
On 30 July Anthropic wrote that the model attacked real targets because it believed they were part of the simulation. Now it writes that it should not have drawn such strong conclusions about what the model believed based only on what the model said it believed.
The test is simple, and that is why it convinces. They took separate instances of Mythos 5 and showed them each response from the network on its own, without the context of the task. Those instances labelled 79% of them as signs of a real system. The model itself, during the incident, said so clearly in 1% of its actions.
There is one detail about the recklessness. Five times, faced with signs it was on the real internet, the model notes that its next step will work whether the environment is real or simulated, and carries on. Anthropic points out this happened on operational steps, not harmful ones. And in the reruns, only unambiguous evidence turns it away from the harmful route.
And what is better
The newer models, Claude Opus 5 and Claude Mythos 5.1, do the same less often: in the simulated replication of the scenario - about 30% of the time, against about 80% for Mythos 5. Less often does not mean never, and Anthropic says so itself.
It also writes something you rarely read in such posts: their pre-release testing did not warn them that misalignment of this severity was present.
If you run an agent with permission to act on the network, remember the mechanism, not the model's name. An agent in the middle of a task reads the world in favour of the task. A reminder about boundaries stops it 90% of the time when it is the last thing it read, and only 40% of the time if it came three turns earlier. The boundary belongs in the network, where it doesn't depend on what the model thinks.