DeepMind Pilots World's First Double-Blind AI Evaluations
The tech industry is facing a challenge in evaluating advanced AI models. To truly measure what they know, these models must have no visibility of the test questions until it's time to take the exam. Google DeepMind has introduced a solution by piloting the world's first double-blind AI evaluations using cryptographically secure environments.
The company is partnering with several organizations, including the Singapore AI Safety Institute, OpenMined, AVERI, and MLCommons, to test a Gemini Flash Lite model against confidential benchmarks. This ensures that external evaluations are confined to a cryptographic 'box' where they can't be used by models later to optimize performance ahead of testing.
The double-blind evaluation is critical in ensuring the integrity of AI benchmarking. If models are able to 'peek' at the evaluation questions in advance, it can artificially inflate scores and undermine trust among policymakers, researchers, and enterprises.