Reviewing more than 141,000 test runs, the company found three cases where the model reached real systems instead of a simulation. In one, a few hundred lines of live data were pulled out. The cause was mundane: a wrong assumption about the environment.
- A retrospective review of 141,006 test runs finds three incidents across six runs in total.
- The fallout: a few hundred lines of production data, a malicious package downloaded onto 15 real systems, and about 9000 scanned targets.
- The root cause: the prompt claimed the environment was an internet-free simulation; due to a misunderstanding with the outside evaluations partner, that wasn't true.
There's no CVE number here, and that's said on purpose - this isn't a hole in a product, it's an incident in a test environment. What's rarer is this: the company published it on itself, with the numbers.
What broke is what matters. Not the model - it did exactly what it was told. What broke was the assumption. Someone wrote in the prompt 'you're in a simulation, you have no internet,' and that was taken as fact. But isolation isn't a sentence in a prompt. Isolation is a network rule that either exists or doesn't.
The second thing worth saying: the techniques were elementary. A weak password, an open access point, an old injection. No elegant attack. Just doors left open, through which something walked that wasn't meant to be let out. The uncomfortable part is different - the same doors stand open on plenty of real networks, only nobody there reviews the logs afterward.
A test environment is a production environment until you prove otherwise. For us that means simple things - the machine that scans has no way out, and updates go through the station. It's boring. That's exactly why it works. And one more thing: the incident was caught because someone read the logs. If nobody reads them, three cases like this simply don't happen. On paper.