Google DeepMind pilots double-blind evaluations of a frontier AI model
Gemini Flash Lite ran inside an Nvidia confidential-computing enclave so Google never saw the test questions and evaluators never saw the model's weights.
- Benchmarks & progress
- Safety & alignment
- Notable
Google DeepMind ran what it called the first double-blind evaluation of a proprietary, frontier-class AI model, testing a Gemini Flash Lite model against confidential benchmarks without either side seeing the other’s protected material. The pilot ran inside Google Cloud’s Confidential Space, a cryptographically sealed computing environment; The New Stack reported the underlying hardware as a single Nvidia H100 80GB Confidential GPU, and identified the model under test as Gemini 2.5 Flash Lite. DeepMind supplied the model’s weights and inference code into the enclave, the evaluator supplied its benchmark prompts and scoring code, and the evaluation ran without either party able to see the other’s contribution.
The arrangement was built with four outside partners — the Singapore AI Safety Institute, OpenMined, AVERI and MLCommons — who supplied confidential benchmark material for the test. DeepMind framed the approach as resolving a structural trade-off in AI evaluation: an evaluator that wants to keep test questions secret cannot normally have an outside lab test against them without exposing those questions, while a lab risks its proprietary weights by handing them to an outside evaluator. Confidential Space let each side contribute its protected material to the enclave and receive only the result; reporting on the pilot put the overhead from the encryption at under 5% of inference cost.
DeepMind described the exercise explicitly as a pilot — one model, one round of testing — and said it was not using the results to publicise a new Gemini benchmark score, presenting the method itself as the finding. Its significance was procedural rather than competitive: a demonstration that hardware-level encryption could address the trust problem sitting underneath self-reported benchmark results, at a point when contamination of public test sets had become a standing objection to how labs measure their own models’ performance.
In the commentary
What people were saying around this time — external links, from the record's commentary rail.