← 返回信息流

arXiv cs.AI论文

准确但不谦逊:评估知识冲突下 LLM 智能体的认知谦逊度

arxiv.org作者:Kaiser Sun, Bernal Jimenez Gutierrez, Hongjun Liu, Jingyu Zhang, Jie Gao, Mark Dredze, Daniel Khashabi论文AI评分:70/100

该研究针对检索证据与智能体先验信念冲突的场景,提出评估“认知谦逊”(Epistemic Humility)的新指标。现有评测多关注任务成功率,忽视了智能体处理冲突时的表现。作者通过实验考察智能体在面临矛盾信息时,是修正答案、承认不确定性还是坚持错误结论,旨在填补对智能体不确定性与自我修正能力的评估空白。

这篇是正式发表的长论文,站内提供中文解读,全文请到原文阅读 PDF。

Abstract:When retrieved evidence contradicts an agent's prior beliefs, does it revise its answer, acknowledge uncertainty, or persist with an incorrect conclusion? Existing evaluations of agentic systems focus primarily on task success, offering limited insight into how agents handle such conflicts. We propose to evaluate agents on epistemic humility (EH): the agent's willingness to recognize, act on, and communicate uncertainty during task execution. We operationalize EH through three trajectory-level behavioral dimensions: Identify, Solve, and Escalate (ISE). Through knowledge conflict, situations where the backbone language model's parametric knowledge contradicts the evidence it encounters, or where two contextual sources disagree, we evaluate two conflict settings: (1) controlled conflict and (2) naturally occurring conflict during multi-step agentic execution, each paired with matched no-conflict controls. Evaluating four agents, we find that higher task accuracy does not necessarily correspond to greater epistemic humility: some high-accuracy configurations recognize conflicts during execution but do not communicate unresolved uncertainty in their incorrect final answers. Trajectory-level analysis further reveals that agents frequently detect conflicts in early steps of execution but fail to maintain or resolve them in later steps. Finally, we show that model-level interventions can improve EH, but often at the cost of task accuracy, suggesting that epistemic humility emerges from the interaction among the backbone model, agent harness, and evaluation environment.

阅读原文