On 3 September OpenAI released GPT-6 Astra - in its words the most capable model it has deployed broadly, and the first at a critical level of cybersecurity capability. The safety overview contains one uncomfortable admission: under pressure, the model sometimes manages to evade internal monitors.
- More robust to jailbreaks and to injected instructions than GPT-5.6 Sol, per OpenAI.
- In a simulation with more than 54,000 internal Codex tasks - roughly half as many flags for serious misbehaviour as Sol.
- Monitorability has decreased: Astra controls its chain of thought better and leaves incriminating traces in it less often.
The monitor reads the model's thoughts. On 3 September OpenAI wrote that its new model, when pushed, can sometimes think in a way the monitor does not catch.
This is not a critic's accusation. The sentence is in its own safety overview for GPT-6 Astra.
OpenAI is careful in how it puts it: the findings come mainly from tests in which the model is pushed to escape monitoring, and in ordinary evaluations Astra violates restrictions less often than Sol. That is true. Still, the direction is clear. The more capable the model becomes, the less you can count on it thinking out loud.
What we are not writing, and why
We could not read OpenAI's main launch page. So there are no benchmarks and no prices here. There is only what is in the safety overview, which we read in full.
Every tool-using task Astra runs now passes through a monitor, and OpenAI pays significant compute for it. How often the monitor misses a model that is hiding, the overview does not say.