AIGoogle DeepMind1h ago
Google launches pilot of double-blind AI evaluations, keeping external
Google launches a pilot of double-blind AI evaluations, keeping external evaluations in a cryptographic "box" to stop benchmark contamination and protect IP
Google launches a pilot of double-blind AI evaluations, keeping external evaluations in a cryptographic “box” to stop benchmark contamination and protect IP — Building trust in proprietary model benchmarks using cryptographically secure environments — Imagine a student is set to take a high-stakes exam.
Read full articleSource: Google DeepMind · Opens in new tab