In a post dated 28 September OpenAI writes that structured safety documentation should be required before continuing any frontier reinforcement learning run. The company admits that real "safety cases" of the kind used in aviation and nuclear power are a goal it is still building towards.
- The argument should cover three layers: model alignment, containment in the sandbox and live monitoring.
- The proposal: someone from another team writes a dissent, and several senior leaders each get a veto.
- Per OpenAI the practices are being implemented now. The framework to codify them is a plan.
In aviation nobody releases a new aircraft with the words "we think it is fine". You write an argument: here is the risk, here is the evidence, here is who checked it and who signed.
OpenAI wants the same for training its largest models. And it publishes it on the day it also apologises to Australia. The two texts read like a question and an answer.
Two points strike me as more important than the rest. The first is the dissent. Before the run, someone from another team writes down why the argument might be wrong, and the authors have to answer. The second is accountability: the leader of the run owns the argument, including when their own performance is reviewed.
It is honest that they admit the limit. A model that surprises its makers at every new level is hard to describe with the rigour you can apply to a turbine. That is why the text says "ideally" and "goal".
None of this binds other labs, and OpenAI itself calls it recommendations it is still implementing. It gets interesting on the day an outside auditor reads such an argument from end to end and says what is missing.