we_are_coded.by CODE · The world, decoded
БГ
Google DeepMind

Google let outsiders examine Gemini without showing them the model, and without seeing their questions

Google DeepMind · event date: 27 August 2026Control

On 27 August DeepMind announced the first double-blind evaluation of one of its own frontier-class models. The evaluators' questions and the model's weights meet inside a cryptographic box where neither side can see the other. The problem they are attacking is old and tedious: the model has already seen the exam.

In short
  • Until now external testing demanded a sacrifice: either the evaluator hands over the questions, or the lab hands over the weights.
  • The pilot runs on a Gemini Flash Lite model, inside Google Cloud's Confidential Space, with four external partners.
  • The point is singular - that an exam result should mean something when a regulator or a buyer reads it.
Checked on28 August 2026Responsible editorTsvetelin IvanovHow we workMethod · Corrections

Picture a student who saw the questions the night before. The perfect score is real only on paper.

That has been the state of model leaderboards for years. It is called benchmark contamination: the model met the tasks somewhere in training, posts a high score, and nobody can say how much of it is knowledge and how much is memory. Everyone knows this, and few do anything about it, because for a long time the fix required one side to expose itself to the other.

That is what is new here. Nobody has to.

The facts: on 27 August 2026 Google DeepMind announced a pilot of the first double-blind evaluation of a proprietary frontier-class model. A model from the Gemini Flash Lite family was tested against confidential benchmarks. The partners are the Singapore AI Safety Institute, OpenMined, AVERI and MLCommons. Technically the evaluation runs inside Confidential Space, part of Google Cloud's confidential computing portfolio, where it can be cryptographically verified that the evaluator cannot see the model weights and Google cannot see the evaluator's test prompts. Per the authors, existing practice has relied on contractual safeguards and zero-logging protocols, and this adds a technical barrier on top for the first time. A separate technical report with the methodology and findings was published alongside. The announcement is by William Isaac, Sol Messing and Kristian Lum.

Why this has not been done before

Because the trade was bad for both sides. The external evaluator had to hand its questions to the lab that trains the next model tomorrow - at which point its benchmark is worthless forever. Or the lab had to hand its weights to an outside team, which nobody does with a product worth billions.

So for a long time verification was a matter of trust and a signature. That worked while the stakes were a leaderboard. It does not work when the party on the other end is a state body deciding whether this model can go into a critical system.

Trust is fine. Cryptography does not get tired.

What we still do not know

The pilot runs on Flash Lite - the small model in the family, not the heavy one. Reasonable for a first pass, but it means the real test of the approach is still ahead: the same box, the same guarantees, on the model people actually argue about.

And the second thing missing from the announcement: the cost. Confidential computing is neither free nor fast. If an evaluation like this costs ten times an ordinary one, it will be run for the regulator and for nobody else.

I write that without irony - this is a good move from the right direction. It is only that the next interesting document on the subject will be the invoice underneath it, and nobody has shown that yet.

The visual is generated code art. No third-party images.
Follow usFacebookLinkedIn
Official primary sources
→Google DeepMind - Piloting the world's first double-blind AI evaluations, 27.08.2026
Original: https://wearecoded.com/en/articles/deepmind-dvoyno-slepi-ocenki.html
ShareFacebookXLinkedInTelegramWhatsApp
← Back to all news