← 返回信息流

Google DeepMind推文

谷歌率先试点前沿AI双盲评估

政策监管AI评分:70/100

谷歌宣布率先试点前沿AI双盲评估,通过创建安全环境,不泄露测试提示和模型权重,确保外部安全与性能评估的私密性、稳健性和可信度。

译文

在行业首创中,我们正在试点对前沿AI进行双盲评估。通过创建一个安全环境,既不公开测试提示词,也不公开模型权重,我们可以确保对我们模型的外部安全性和性能评估保持私密、稳健且可信。→ https://goo.gle/3St2xan

Google DeepMind

@GoogleDeepMind

In an industry first, we’re piloting double-blind evaluations for frontier AI. By creating a secure environment where neither test prompts nor model weights are revealed, we can ensure external safety and performance evaluations of our models remain private, robust, and trustworthy. → https://goo.gle/3St2xan

阅读原文