Deepmind Tests Gemini Inside Sealed Evaluation System Arabian Post
The pilot represents a significant attempt to address benchmark contamination, one of the most persistent weaknesses in measuring advanced AI systems. Under conventional testing arrangements, model developers may gain access to evaluation questions, while independent testers can sometimes require access to proprietary model weights. Either situation can compromise commercially sensitive information or weaken confidence that a model is being assessed on material it has never encountered.
DeepMind worked with the Singapore AI Safety Institute, independent evaluation organisation AVERI, privacy technology group OpenMined and benchmarking consortium MLCommons. The experiment placed Gemini 2.5 Flash-Lite and confidential evaluation material inside a protected computing environment, allowing the model to process the questions without exposing them to Google and without revealing the model weights to the evaluators.
Legal Disclaimer:
MENAFN provides the
information “as is” without warranty of any kind. We do not accept any
responsibility or liability for the accuracy, content, images, videos,
licenses, completeness, legality, or reliability of the information
contained in this article. If you have any complaints or copyright issues
related to this article, kindly contact the provider above.

Comments
No comment