← 返回信息流

DAIR.AI推文

任务模型归纳:从被动记录中提取可复用的智能体技能

论文AI评分:70/100

本文提出任务模型归纳(TMI)方法,从用户工作的原始屏幕录制和鼠标键盘事件中,自动发现潜在任务并构建符号化任务模型。TMI 首先将多线程轨迹分离为独立任务,与人工标注分组的一致性达 0.974;随后为每个任务生成层次化目标模型和控制流程序模型,可重建 74.9% 的执行步骤。基于这些任务模型提取的技能,在保留任务上的准确率比最强基线提升 30.0%。该方法利用现有工作电脑上的被动记录,将其转化为可审计、可复用的技能。

译文

这是提升智能体工作流最有效的方法之一。如果你正在构建计算机操作型智能体,这个方法值得你花时间了解。任务模型归纳(Task Model Induction)会获取某人工作的原始录制内容——仅包含截图、鼠标和键盘事件——并将其转化为关于工作实际完成方式的符号化模型。难点在于,真实录制内容是多线程的。人们会在任务中途切换目标。TMI首先在无约束的轨迹中发现潜在任务并将其分离,与真实分组相比,一致性达到0.974。每个恢复出的任务随后会获得两样东西:一个关于目标如何分解的层级目标模型,以及一个组织执行过程的控制流程序模型。它能重建74.9%的已观察执行步骤,而基于这些任务模型提炼出的技能,在保留任务上的准确率比最强的工作流归纳基线高出30.0%。被动轨迹数据已经存在于大多数工作笔记本电脑上。这项研究只是展示了如何将它们挖掘为可审计、可复用的技能。论文:https://arxiv.org/abs/2608.20319 在我们的学院中追踪更多热门AI论文:https://academy.dair.ai/

DAIR.AI

@dair_ai

This is one of the most effective ways to improve your agentic workflows. If you are building computer-use agents, this one is worth your time. Task Model Induction takes a raw recording of someone working, just screenshots and mouse and keyboard events, and turns it into a symbolic model of how the work was actually done. The hard part is that real recordings are multi-threaded. People switch between goals mid-task. TMI first discovers the latent tasks inside an unconstrained trace and separates them, hitting 0.974 agreement against ground-truth groupings. Each recovered task then gets two things. A hierarchical objective model of how the goal decomposes, and a procedure model of the control flow that organized execution. It reconstructs 74.9% of observed execution steps, and skills derived from these task models lift held-out task accuracy by 30.0% over the strongest workflow induction baseline. Passive traces are sitting on most work laptops already. This work just shows how to mine them into auditable, reusable skills. Paper: https://arxiv.org/abs/2608.20319 Track more trending AI papers in our academy: https://academy.dair.ai/

阅读原文