Google DeepMind Seals Gemini Test to Protect Benchmarks
AI benchmarks are supposed to reveal what models can do, but Google DeepMind is now putting the tests behind a cryptographic wall to make sure the models have not seen the answers first. Google DeepMind said Thursday that it has piloted what it describes as the first double-blind evaluation of a proprietary frontier AI model,…
