Anime, manga, and games, with a take · A Yukimedia publication

← all stories other 1 sources · 50m ago ·

Google DeepMind Pilots Double-Blind AI Evaluations

The pilot aims to protect both test questions and model weights, addressing a core reliability problem where benchmark scores can be inflated by cheating, and the group behind the pilot report has published a report on the first double-blind evaluation of a proprietary language model.

Reporting from 1 source: GIGAZINE.

Google DeepMind Pilots Double-Blind AI Evaluations

Google DeepMind has devised a testing method that evaluates AI models while preventing cheating by the models themselves and by the companies that develop them. The method uses Confidential Computing, a security system within Google Cloud virtual machines that encrypts data during processing. Both external evaluation institutions and development companies send their data to the secure space for testing.

AI benchmarks only work if the scores are honest. Google DeepMind's new approach targets the two ways those scores get corrupted: a model seeing test questions in advance, or a developer training a model to game a specific benchmark.

The method runs evaluations inside a confidential space in Google Cloud virtual machines, a security system that encrypts data while it is being processed. The external evaluation institution sends its test questions, and the development company sends its model weights, so neither side has to risk leaking its data to the other.

Google DeepMind, the Singapore AI Institute, OpenMined, MLCommons, and AVERI already ran an initial pilot using MLCommons' safety benchmark family with Gemini 2.5 Flash-Lite. The group published a report on the first double-blind evaluation of a proprietary language model through AVERI.

Synthesized by Yomimono from the 1 cited source below, including Japanese-language reporting where cited, then editorially reviewed before publishing.

Sources