← 返回信息流

DAIR.AI推文

微软关于智能体在真实业务流程中可靠性的重磅论文

论文评测AI评分:90/100

微软提出Thinkingbox,一个用于评估AI智能体在真实业务流程中可靠性的沙箱环境,包含507个跨零售、酒店、汽车保险、新银行IT和咨询支持的政策条件工作流基准。每个尝试根据智能体留下的后端状态进行评分,可执行检查接受有效轨迹并拒绝错误、缺失或额外效果。最强模型达到65.36%的pass@1和25.25%的pass^20,许多失败尝试以有效的状态更改工具调用干净终止,仅观察响应或工具调用难以判断任务是否真正完成。

译文

微软的一篇重磅论文,主题是智能体在真实业务流程中的可靠性。(建议收藏)Thinkingbox 是一个沙盒环境,提供隔离的、兼容 MCP 的工具会话,并附带一个包含 507 个策略条件化工作流的基准测试集,覆盖零售、酒店、汽车保险、新银行 IT 以及咨询支持等领域。每次尝试都会根据智能体在后台留下的最终状态进行评分。可执行检查会接受有效的轨迹,并拒绝错误、缺失或多余的影响,因此附带损害也会被计入扣分。目前最强的模型达到了 65.36% 的 pass@1 和 25.25% 的 pass^20。许多失败的尝试都以有效的状态变更工具调用干净地终止了。仅观察响应或工具调用,几乎无法判断任务是否真正完成。论文链接:https://arxiv.org/abs/2608.19741 在我们的学院中追踪更多热门 AI 论文:https://academy.dair.ai/

DAIR.AI

@dair_ai

Banger paper from Microsoft. It's on agent reliability in real business workflows. (bookmark it) Thinkingbox is a sandbox with isolated MCP-compatible tool sessions, plus a benchmark of 507 policy-conditioned workflows across retail, hospitality, auto insurance, neobank IT, and consulting support. Every attempt is graded on the backend state the agent leaves behind. Executable checks accept valid trajectories and reject wrong, missing, or extra effects, so collateral damage counts against you. The strongest model reaches 65.36% pass@1 and 25.25% pass^20. Many failed trials terminate cleanly with valid state-changing tool calls. Watching the response or the tool call tells you very little about whether the task actually completed. Paper: https://arxiv.org/abs/2608.19741 Track more trending AI papers in our academy: https://academy.dair.ai/

阅读原文