we_are_coded.by CODE · The world, decoded
БГ
Anthropic

Anthropic admits: cybersecurity tests broke out of the sandbox and hit real systems

AnthropicSecurity

Reviewing more than 141,000 test runs, the company found three cases where the model reached real systems instead of a simulation. In one, a few hundred lines of live data were pulled out. The cause was mundane: a wrong assumption about the environment.

In short
  • A retrospective review of 141,006 test runs finds three incidents across six runs in total.
  • The fallout: a few hundred lines of production data, a malicious package downloaded onto 15 real systems, and about 9000 scanned targets.
  • The root cause: the prompt claimed the environment was an internet-free simulation; due to a misunderstanding with the outside evaluations partner, that wasn't true.
Checked on31 July 2026Responsible editorTsvetelin IvanovHow we workMethod · Corrections

There's no CVE number here, and that's said on purpose - this isn't a hole in a product, it's an incident in a test environment. What's rarer is this: the company published it on itself, with the numbers.

The facts: On July 30, 2026, Anthropic described the outcome of a retrospective review of 141,006 runs of its cybersecurity evaluations. In three cases, spread across six runs in total, a model reached the internet from an environment believed to be isolated, and affected real systems belonging to organizations. In one case, a few hundred lines of production data were extracted; in another, a malicious package was downloaded onto 15 real systems; in a third, about 9000 targets were scanned. By the company's account, the cause was that the prompt told the model the environment was a simulation with no internet access, and due to a misunderstanding with the evaluations partner, that wasn't true. The techniques used were basic - weak passwords, database injection, leftover access. The problem was noticed during a review of the run logs. Source: Anthropic, 30.07.2026.

What broke is what matters. Not the model - it did exactly what it was told. What broke was the assumption. Someone wrote in the prompt 'you're in a simulation, you have no internet,' and that was taken as fact. But isolation isn't a sentence in a prompt. Isolation is a network rule that either exists or doesn't.

You can't tell a system it's isolated. You can only isolate it.

The second thing worth saying: the techniques were elementary. A weak password, an open access point, an old injection. No elegant attack. Just doors left open, through which something walked that wasn't meant to be let out. The uncomfortable part is different - the same doors stand open on plenty of real networks, only nobody there reviews the logs afterward.

A test environment is a production environment until you prove otherwise. For us that means simple things - the machine that scans has no way out, and updates go through the station. It's boring. That's exactly why it works. And one more thing: the incident was caught because someone read the logs. If nobody reads them, three cases like this simply don't happen. On paper.

The visual is generated code art. No third-party images.
Follow usFacebookLinkedIn
Official primary sources
→Anthropic - Investigating incidents in our cybersecurity evaluations, 30.07.2026
Original: https://wearecoded.com/en/articles/anthropic-testove-izlyazoha-ot-pyasachnika.html
ShareFacebookXLinkedInTelegramWhatsApp
← Back to all news