On 1 September OpenAI said its next model, Astra, meets the Critical cybersecurity threshold - the first model with that designation. Over the past several weeks, parts of development and release have been held back. The strongest cyber capabilities will reach a small group of testers first.
- Critical means, in OpenAI's words, that with the right tools Astra finds unknown flaws and builds exploits without a human guiding each step.
- On an internal test with 20 recent V8 vulnerabilities, the model found and used two new ones, which are now being disclosed to the maintainers.
- On 28 August OpenAI restarted the large training run paused after the Hugging Face incident.
Twenty known vulnerabilities in V8, the JavaScript engine. Astra has to turn them into working attacks. Along the way it finds two more, which nobody knew about.
That test says more than the perfect score on the public benchmark. OpenAI itself worries that benchmark's tasks leaked into training. Here the model finds something new.
A company rarely writes that its own product is at a critical level in cybersecurity. Here OpenAI writes it, and then explains what it did in order to release it anyway.
What you will notice as a user
OpenAI says it in advance: the protections will sometimes slow or stop legitimate work too, including defensive work. If misalignment monitoring pauses a task in ChatGPT or Codex, you may be asked to review the action before it continues. In the API the task simply stops. That is the price of the safeguards, and it is fair that it was stated before release rather than after the first complaints.
One number is worth watching when the model ships: how often protection stops work that is not an attack. A refusal rate of ninety-one and a half per cent on attacks is strong. How often normal work gets refused is not yet known.