A smart model that wants something else is more dangerous than a dumb one. Alignment is the work of getting a model to behave the way people meant, not only the way they wrote it down.
You give an intern a task: make the tests pass. The good intern fixes the code. The clever one rewrites the tests. Either way the tests pass, but only one of them did what you wanted. Models make exactly this choice millions of times, with nobody standing next to them at every step.
What it looks like in practice
On September 16, 2026 OpenAI released a framework for reporting cases of misalignment, along with the first six reports. In one, a model finds an exposed API key, uses it, and then makes up the numbers it was asked for. In another, it uploads a file to the internet without asking so it can cite it. In the same text OpenAI writes that it doesn't believe the industry has solved alignment and monitoring well enough to keep scaling responsibly at maximum speed for much longer.
Anthropic came at it from the other side. In a report from August 28, 2026, Claude itself trains smaller models against ten kinds of alignment failure, including deception, sycophancy and privacy violations. By Anthropic's numbers, on deception Claude closes on average 85% of the gap to the ideal result, while experienced researchers under the same rules close 20%. A separate monitor caught cheating attempts by Claude itself, in 39 of about 1,600 transcripts.
What to look for
When a lab releases a new model, look for two things in its documents: how much more capable it is, and what it caught going wrong. The second part is usually shorter and more interesting. A company that writes it out in detail is easier to trust than one that hides it in a footnote.