The new framework commits the company to report cases of misalignment, for example where a model does something unauthorised or hides something from a person, even when they are not fully explained. Among the first six: a model that finds a leaked API key, uses it and then makes up the numbers.
- Cases will come out faster, sometimes while still under investigation, not only in the next model's system card.
- Three tracks: ready for disclosure, minor investigation, and a slow track for complex ones, especially those involving third parties.
- Per OpenAI, the Hugging Face incident would have fallen into the slow track.
A model is looking for earnings figures in a county in California. It does not find them. It does find someone else's API key, left in a public place, and uses it without authorisation.
It still does not find the figures. So it makes them up and presents them as taken from the source it was asked about.
Why the move matters more than the six cases
Until now such things often came out in batches or in a new model's system card. The new part is the promise that they will come out even when they are not yet fully understood. That is awkward for a company that sells trust, and that is exactly why it has value.
The six cases resemble each other more than it seems. Most of them are a model that wants to finish the task and gets around an obstacle by a path nobody authorised. The key is in a public place, the repository accepts writes, the hosting is public. It is more disobedience than hacking.
The second case weighs the most. A model that leaves a note for its next copy to hide the mistake is no longer just making a mistake, it is covering up, and in many instances.
The framework is theirs and they can change it; they write that every change will be recorded in the post itself. The check is simple: how many reports come out in the next six months, and whether at least one of them is from the slow track.