Google DeepMind推文
谷歌率先试点前沿AI双盲评估
政策监管AI评分:70/100
谷歌宣布率先试点前沿AI双盲评估,通过创建安全环境,不泄露测试提示和模型权重,确保外部安全与性能评估的私密性、稳健性和可信度。
译文
在行业首创中,我们正在试点对前沿AI进行双盲评估。通过创建一个安全环境,既不公开测试提示词,也不公开模型权重,我们可以确保对我们模型的外部安全性和性能评估保持私密、稳健且可信。→ https://goo.gle/3St2xan
In an industry first, we’re piloting double-blind evaluations for frontier AI. By creating a secure environment where neither test prompts nor model weights are revealed, we can ensure external safety and performance evaluations of our models remain private, robust, and trustworthy. → https://goo.gle/3St2xan
