The model that thinks first and answers after. Its thinking shows up as text, so researchers read it like a diary. The question is how long the diary will keep telling the truth.
Remember solving problems at school. The teacher wanted not just the answer but the working on the page. Not because your handwriting was interesting, but because the working shows where you went wrong and whether you copied. The chain of thought is exactly that working on the page.
The diary that might stop telling the truth
More than 40 researchers, among them Yoshua Bengio and OpenAI's Jakub Pachocki, wrote in a July 2025 paper that the chain of thought offers a rare chance: to read a model's intent to misbehave before it acts. In the same text they warn that this opportunity may turn out to be fragile.
On September 3, 2026 OpenAI released GPT-6 Astra and wrote itself that the model is harder to monitor than its predecessor. It controls its own chain of thought better and is less likely to put incriminating information in it. When pushed on purpose to evade the monitors, it sometimes succeeds. The company found no signs of hidden encoded reasoning.
Why the draft is hidden from the grader
When you read that a lab hides the chain of thought from its automated graders, that isn't red tape. OpenAI says it plainly in its proposals from September 28, 2026: if the grader sees the draft, the model can learn to write drafts that slip past monitoring. You keep the diary honest by not punishing what's written in it.