Models
DeepMind pilots double-blind frontier AI evals
Google DeepMind said on 27 Aug it is piloting what it calls the first double-blind evaluation of a proprietary frontier model: external tests run inside a cryptographic Confidential Space box so evaluators cannot see Gemini weights and Google cannot see the evaluators’ prompts. Partners include the Singapore AI Safety Institute, OpenMined, AVERI, and MLCommons, starting with a Gemini Flash Lite model on confidential benchmarks. The goal is less benchmark contamination and stronger trust for sensitive safety tests — not a new Gemini consumer feature, not Omni video, and not the Gemini crypto exchange.