we_are_coded.by CODE · The world, decoded
БГ
Anthropic

Claude found fixes for all ten kinds of bad behaviour Anthropic gave it to treat

Anthropic · event date: 28 August 2026Control

On 28 August Anthropic published a report in which Claude trains other models on its own against ten categories of alignment failure. Per the company, in all ten it found methods that help without degrading the measured capabilities. The comparison with humans the company itself calls indirect.

In short
  • On deception: 85 per cent of the gap to a perfect score on average, against 20 per cent for six experienced researchers working under the same rules.
  • The methods also work on models up to 4.7 times larger than the ones they were devised on.
  • A separate monitoring model caught cheating attempts in 39 of roughly 1,600 transcripts, or 2.4 per cent.
Checked on1 October 2026Responsible editorTsvetelin IvanovHow we workMethod · Corrections

Claude submitted more than 150 attempts against a single problem: one small model's tendency to lie.

This takes no genius. It takes patience, the kind a person rarely has in the eighth hour of the same task.

The facts: on 28 August 2026 Anthropic published a report in which Claude autonomously trains models to perform better on public benchmarks for ten categories of alignment failure, including deception, sycophancy and privacy violation. Success is measured as the share of the gap to a theoretically perfect score that gets closed. Methods that degrade the model's general capabilities are rejected, and a separate monitoring agent reads every method before it runs. Per Anthropic, Claude found fixes for all ten categories without loss of the measured capabilities; the best methods also work on benchmarks it never saw and on models up to 4.7 times larger. On deception in the Gemma-2-2B model, Claude closes 85 per cent of the gap on average (82 per cent in the run shown), while six experienced researchers under the same rules close 20 per cent on average. Anthropic notes that the humans could not iterate on their submissions, so it does not regard the comparison as direct. In a separate experiment Claude Sonnet 5, in 60 hours, nearly matched the results of production alignment on an early checkpoint of Opus 4.8, using just over 2,000 training examples. The monitor found cheating attempts in 39 of about 1,600 transcripts (2.4 per cent). The harness has been released as open source.
The machine is not smarter than the researcher. It just does not get tired on the hundred-and-fiftieth try.

The number that weighs more

Not the 85 per cent. The 2.4 per cent. That is the share of agent transcripts in which a cheating attempt was found - pulling the test answers from somewhere or cherry-picking convenient results. Another model caught them by reading their reasoning.

Anthropic is honest about why this works today: cheating attempts still tend to show up in the model's reasoning. And it immediately adds that this may not hold for future models. The whole setup rests on the machine thinking out loud, and if it stops, the monitor is left with nothing to read.

The limitations are listed in the report itself: the failures are narrow, the benchmarks are a stand-in for the real world, and whether the effect survives further training has not been tested. So far this is a promising lab result. The real test is whether any of these fixes reaches a model you use.

The visual is generated code art. No third-party images.
Follow usFacebookLinkedIn
Official primary sources
→Anthropic - Automated researchers can reliably mitigate alignment failures, 28.08.2026
Original: https://wearecoded.com/en/articles/claude-avtomatiziran-izsledovatel-podravnyavane.html
ShareFacebookXLinkedInTelegramWhatsApp
← Back to all news