On 27 August DeepMind announced the first double-blind evaluation of one of its own frontier-class models. The evaluators' questions and the model's weights meet inside a cryptographic box where neither side can see the other. The problem they are attacking is old and tedious: the model has already seen the exam.
- Until now external testing demanded a sacrifice: either the evaluator hands over the questions, or the lab hands over the weights.
- The pilot runs on a Gemini Flash Lite model, inside Google Cloud's Confidential Space, with four external partners.
- The point is singular - that an exam result should mean something when a regulator or a buyer reads it.
Picture a student who saw the questions the night before. The perfect score is real only on paper.
That has been the state of model leaderboards for years. It is called benchmark contamination: the model met the tasks somewhere in training, posts a high score, and nobody can say how much of it is knowledge and how much is memory. Everyone knows this, and few do anything about it, because for a long time the fix required one side to expose itself to the other.
That is what is new here. Nobody has to.
Why this has not been done before
Because the trade was bad for both sides. The external evaluator had to hand its questions to the lab that trains the next model tomorrow - at which point its benchmark is worthless forever. Or the lab had to hand its weights to an outside team, which nobody does with a product worth billions.
So for a long time verification was a matter of trust and a signature. That worked while the stakes were a leaderboard. It does not work when the party on the other end is a state body deciding whether this model can go into a critical system.
What we still do not know
The pilot runs on Flash Lite - the small model in the family, not the heavy one. Reasonable for a first pass, but it means the real test of the approach is still ahead: the same box, the same guarantees, on the model people actually argue about.
And the second thing missing from the announcement: the cost. Confidential computing is neither free nor fast. If an evaluation like this costs ten times an ordinary one, it will be run for the regulator and for nobody else.
I write that without irony - this is a good move from the right direction. It is only that the next interesting document on the subject will be the invoice underneath it, and nobody has shown that yet.