A Test Neither Side Could Read
On August 27, 2026, Google DeepMind ran what it calls the world's first double-blind evaluation of a proprietary, frontier-class AI model, and neither party in the test could see the other's most sensitive material. The evaluation used Confidential Space, part of Google Cloud's Confidential Computing portfolio, to run the check inside a verified cryptographic enclave: the evaluator's test prompts stayed encrypted and out of Google's reach, and Google's model weights stayed encrypted and out of the evaluator's reach.
"By using Confidential Space within Google Cloud's Confidential Computing portfolio, we can cryptographically verify that both the external evaluation data and the proprietary model remain private to their respective owners," DeepMind wrote in its announcement. Four partners joined the pilot as evaluators: the Singapore AI Safety Institute, the privacy-technology nonprofit OpenMined, the assurance group AVERI, and the standards body MLCommons.
That arrangement removes a tradeoff that has shaped every AI evaluation until now. An outside evaluator either has to trust a lab's self-reported scores, or the lab has to hand over model weights it treats as its core trade secret, or the evaluator has to hand over its test questions and risk the model simply memorizing the answer key for next time. This pilot is the first public case of a proprietary, frontier-class model going through a real test without either side making that trade.
Why Benchmark Scores Have Been Easy To Game
The problem this pilot is built to fix has a name: benchmark contamination, the situation where a model has already been exposed, directly or indirectly, to the questions it is later scored on. When that happens, a benchmark stops measuring what a model can actually do and starts measuring what it has already memorized.
The scale of the problem is not small. Earlier research cited in coverage of the pilot found signs of benchmark leakage in about half of 31 AI models tested, with contamination inflating scores more for larger models than smaller ones. That matters because benchmark scores are not just marketing copy; labs also use them as evidence for safety claims, and regulators and enterprise buyers increasingly treat them as compliance signals rather than treating them with real skepticism.
The Model Being Tested Is Not The One That Matters Most
Gemini Flash Lite, the model DeepMind put through this pilot, sits at the smallest and cheapest tier of the Gemini family, not the frontier flagship whose capability and safety claims carry the most commercial and regulatory weight. Testing a lightweight model first is a reasonable way to prove out new plumbing, but it also means the pilot has not yet answered the question that matters most: whether the same cryptographic check works, and gets run, on the model Google actually sells as its frontier system.
The four evaluator partners were chosen by Google itself, and the Singapore AI Safety Institute is a genuine state safety body, which gives the pilot some real independence. But no outside auditor has yet verified the integrity of the enclave itself, and DeepMind has described this as a pilot, not a standing commitment to test every model this way going forward.
What This Means For AI Buyers And Regulators In The EU And UK
For an EU or UK organization buying or governing AI systems, the interesting question is not whether this particular pilot worked, but whether double-blind evaluation becomes the standard proof-of-evaluation that vendors are expected to produce. Independent verification of a frontier model's safety claims today usually means one of two unsatisfying options: trust the vendor's own published model card, or demand weight access that almost no lab will grant.
This pilot is a working technical answer to that gap, not yet a legal requirement anywhere, and not yet applied to a model whose claims actually drive procurement decisions or regulatory filings. An EU or UK buyer evaluating a frontier AI vendor for compliance evidence has a concrete new question to ask: will you submit your flagship model, not just your cheapest tier, to an evaluation like this one.
Read next: Qwen's 3 Billion Downloads Hide a Compliance Gap | Anthropic's Wellbeing Grants Repeat OpenAI's Gap



