← 返回信息流

arXiv cs.AI论文

先验还是反馈?大语言模型在适配神经算子时利用什么信息

arxiv.org作者:Julian Chan, Javier Mora Jimenez论文AI评分:70/100

该研究探讨大语言模型(LLM)科学智能体在神经算子适配任务中,是仅依赖初始任务上下文,还是根据实验反馈调整决策。研究设定LLM在有限试验预算下选择微调配置,结果显示其在偏微分方程(PDE)族内及跨族的迁移中,其测试nRMSE低于随机搜索和贝叶斯优化。

这篇是正式发表的长论文,站内提供中文解读,全文请到原文阅读 PDF。

Abstract:Do LLM scientific agents rely only on their initial task context, or do they adapt their decisions in response to experimental feedback? We study this question in neural operator adaptation, where a large language model (LLM) selects fine-tuning configurations under a limited trial budget. Across transfers within and between partial differential equation (PDE) families, the LLM achieves lower held-out test nRMSE than random search and Bayesian optimisation in nearly every matched comparison. Endpoint performance alone cannot distinguish what happens, so we verify each attribution with controlled interventions. Before observing any validation score, the LLM's first configuration already ranks near the top of the corresponding random-search pool, indicating a useful initial bias. A complementary cold-start intervention shows that the selected base learning rate shifts with the PDE description. Once feedback becomes available, reassigning validation scores among evaluated configurations changes the next proposal in every case tested, whereas a value-preserving rewrite produces no comparable aggregate effect. These interventions establish that the LLM's decision-level actions respond to the given task and observed outcomes, showing that it combines a task-dependent prior with sensitivity to experimental feedback.

阅读原文