Tuesday, 02 January 2024 12:17 GMT

Deepmind Tests Gemini Inside Sealed Evaluation System Arabian Post


(MENAFN- The Arabian Post) clearfix"> Google DeepMind has completed what it describes as the world's first double-blind evaluation of a proprietary frontier-class artificial intelligence model, testing Gemini 2.5 Flash-Lite while preventing both the developer and external evaluators from accessing each other's confidential material.

The pilot represents a significant attempt to address benchmark contamination, one of the most persistent weaknesses in measuring advanced AI systems. Under conventional testing arrangements, model developers may gain access to evaluation questions, while independent testers can sometimes require access to proprietary model weights. Either situation can compromise commercially sensitive information or weaken confidence that a model is being assessed on material it has never encountered.

DeepMind worked with the Singapore AI Safety Institute, independent evaluation organisation AVERI, privacy technology group OpenMined and benchmarking consortium MLCommons. The experiment placed Gemini 2.5 Flash-Lite and confidential evaluation material inside a protected computing environment, allowing the model to process the questions without exposing them to Google and without revealing the model weights to the evaluators.

MENAFN29082026000152002308ID1111595481



The Arabian Post

Legal Disclaimer:
MENAFN provides the information “as is” without warranty of any kind. We do not accept any responsibility or liability for the accuracy, content, images, videos, licenses, completeness, legality, or reliability of the information contained in this article. If you have any complaints or copyright issues related to this article, kindly contact the provider above.



More Story