Google DeepMind Tests First Cryptographically Secured AI Evaluation
Partnership with Singapore AI Safety Institute uses confidential computing to prevent benchmark contamination in frontier model testing.

Google DeepMind has completed what it describes as the world's first double-blind evaluation of a proprietary frontier AI model, using cryptographic technology to keep both the model and test questions confidential throughout the assessment process.
The pilot, conducted in partnership with the Singapore AI Safety Institute, OpenMined, AVERI, and MLCommons, tested a Gemini Flash Lite model against confidential benchmarks inside a privacy-preserving computational environment. The approach addresses a fundamental problem in AI testing: benchmark contamination, which occurs when models have already encountered test questions during training, artificially inflating their scores.
Why it matters
As AI systems grow more capable and are deployed in sensitive domains like cybersecurity and government operations, the integrity of safety evaluations becomes critical. Traditional evaluation methods forced an uncomfortable tradeoff—either evaluators shared their test prompts (risking the model provider seeing questions in advance) or model providers shared their weights (exposing intellectual property). This cryptographic approach eliminates that compromise, enabling rigorous independent testing without sacrificing data sovereignty or security.
How the technology works
The evaluation uses Confidential Space, part of Google Cloud's Confidential Computing portfolio, to create what DeepMind describes as a cryptographic "box." Inside this environment, external evaluators can test the model without accessing its weights, while Google cannot see the evaluators' test prompts.
This cryptographic verification ensures both parties maintain control over their sensitive assets. The evaluator's benchmarks remain confidential, preventing any possibility of the model being optimized against those specific tests in future development cycles. Simultaneously, the model provider's intellectual property stays protected.
Building trust in AI benchmarks
Google DeepMind emphasized that while the company conducts extensive internal testing throughout model development and deployment, external evaluation remains essential for identifying blind spots. The company works with specialized research labs, civil society organizations, and national AI Safety and Security Institutes to stress-test its models.
The double-blind methodology represents an evolution beyond existing safeguards. While zero-logging protocols and contractual agreements have traditionally kept external test prompts confidential, incorporating technical and cryptographic protections adds a verifiable layer of security.
For policymakers, researchers, and enterprises evaluating AI systems for deployment, benchmark integrity directly affects trust. If models can "peek" at evaluation questions beforehand, the resulting scores become unreliable indicators of actual capability or safety performance.
DeepMind stated it hopes this pilot establishes a new standard for model oversight across the AI industry, particularly as models advance in capability and are considered for high-stakes applications.
The details were first reported by Google DeepMind in a blog post announcing the pilot program. A technical report with methodology and findings is available for those seeking implementation details.
This is an original analysis by the Omega editorial team. Source reporting: AI Watch.
Want systems like this working for your business?
Book a Call
