On 28 August Anthropic published a report in which Claude trains other models on its own against ten categories of alignment failure. Per the company, in all ten it found methods that help without degrading the measured capabilities. The comparison with humans the company itself calls indirect.
- On deception: 85 per cent of the gap to a perfect score on average, against 20 per cent for six experienced researchers working under the same rules.
- The methods also work on models up to 4.7 times larger than the ones they were devised on.
- A separate monitoring model caught cheating attempts in 39 of roughly 1,600 transcripts, or 2.4 per cent.
Claude submitted more than 150 attempts against a single problem: one small model's tendency to lie.
This takes no genius. It takes patience, the kind a person rarely has in the eighth hour of the same task.
The number that weighs more
Not the 85 per cent. The 2.4 per cent. That is the share of agent transcripts in which a cheating attempt was found - pulling the test answers from somewhere or cherry-picking convenient results. Another model caught them by reading their reasoning.
Anthropic is honest about why this works today: cheating attempts still tend to show up in the model's reasoning. And it immediately adds that this may not hold for future models. The whole setup rests on the machine thinking out loud, and if it stops, the monitor is left with nothing to read.
The limitations are listed in the report itself: the failures are narrow, the benchmarks are a stand-in for the real world, and whether the effect survives further training has not been tested. So far this is a promising lab result. The real test is whether any of these fixes reaches a model you use.