Field note
Double-blind model evaluation makes benchmark integrity a property of the execution environment rather than a promise between organizations.
The pilot is timely because benchmark contamination can inflate scores while external evaluators and model providers both have legitimate reasons to protect sensitive assets. Google DeepMind is piloting a proprietary Gemini Flash Lite evaluation with the Singapore AI Safety Institute, OpenMined, AVERI, and MLCommons. The design uses Google Cloud Confidential Space so the evaluator cannot inspect model weights and Google cannot inspect confidential evaluation prompts.
A cryptographically verifiable confidential-computing environment binds the model and hidden benchmark at execution time without handing either party the other's protected material. High-stakes evaluators should treat isolation, attestation, logging, and prompt custody as part of the benchmark contract, not as administrative detail.
The pilot demonstrates an architecture for one model and partner group; it does not yet establish broad interoperability or eliminate flaws in benchmark design and grading. Revise the conclusion after independent replications report attestation details, failure handling, cost, and whether the method prevents operational leakage across multiple providers.
