← 返回信息流

Meta AI推文

Meta 预览 WildArtifactBench 评估框架

评测模型AI评分:70/100

Meta 今日预览 WildArtifactBench,一个内部评估框架,用于评估多模态智能体在复杂真实任务中的表现。它采用胜率和 Elo 评分,由人类和智能体偏好裁判评判,而非严格的标准答案,从而扩展了实际多模态工作流的任务覆盖。Meta 已发布其中 10 个任务。

译文

今天我们还将预览WildArtifactBench,这是一个内部评估框架,旨在评估智能体在多种交付格式下完成复杂真实世界任务的表现。通过使用人类和智能体偏好评判者的胜率与Elo评分,而非严格的标准答案评分细则,它扩展了跨实际多模态工作流的任务覆盖范围。我们从WildArtifactBench中发布10项任务,这标志着我们在衡量多模态智能体实际效用方面迈出了一步:https://research.meta.ai/wild-artifact-bench

M
Meta AI

@AIatMeta

Today we’re also previewing WildArtifactBench, an internal evaluation framework designed to assess agents on complex, real-world tasks across diverse deliverable formats. By using win rates and Elo scores from human and agentic preference judges rather than strict ground-truth rubrics, it expands task coverage across practical multimodal workflows. We’re releasing 10 tasks from WildArtifactBench as a step forward in our ability to measure the real practical utility delivered by multimodal agents: https://research.meta.ai/wild-artifact-bench

阅读原文