AIThe Decoder1h ago

AI benchmarks have a trust problem and Google wants to fix it

AI benchmarks have a trust problem and Google wants to fix it

TL;DRGoogle is testing secret AI evaluations to prevent cheating on benchmark tests.

Why it matters: Transparent, trustworthy AI benchmarks are essential for comparing models and holding companies accountable.

Google Deepmind is testing a double-blind evaluation of a frontier AI model for the first time. Cryptographic protection through Confidential Space is meant to keep Google from seeing the test questions and keep evaluators from seeing the model weights. The pilot project with…

Read full article

Source: The Decoder · Opens in new tab