← 返回信息流

Ethan Mollick推文

关于AI影响的研究:使用旧模型并非天生有害,但需谨慎讨论

观点论文AI评分:50/100

作者认为,使用旧模型(如GPT-4)研究AI影响并非天生有害,但需谨慎讨论。正面能力声明(如AI达到人类水平)具有持久性,而负面能力声明(如AI在某任务上失败)则需更小心,因为未来模型可能改进。作者提出几种可行框架:明确限定模型版本、测量趋势、论证固有局限、关注调节因素或聚焦人类反应。

译文

仅基于旧模型发布关于AI影响的研究并非 inherently 不好,但这需要非常谨慎的讨论,并且必须对非技术读者非常清晰。如果你展示AI能做某事,那没问题。一般来说,一旦AI获得了一项能力,它不会在后续世代中退化。因此,那些展示AI已跨越某个门槛(如“与人类一样好”)或试图确立某种最低影响或效果的论文,即使使用了GPT-4,也仍然成立。如果研究发现是AI在某方面表现不佳或存在偏见,你就需要更加小心。你不能因为GPT-4在某事上失败,就声称AI在该任务上表现不佳,因为当前或未来的模型可能能做到。那么你能声称什么呢?以下只是一些例子:1)你可以直接明确地说:“GPT-4无法完成X。”这是一种很少见的表述方式,因为在很多情况下这并不构成一篇有趣的、可发表的论文。然而,如果你提供了复现工作所需的信息,它可以成为一个有用的基准来衡量进展。2)你可以测量趋势:比较GPT-4、GPT-5和GPT-5.6 Sol或任何其他模型(至少包含一个推理模型非常重要)。然后你可以就与该任务相关的能力做出更有力的声明。3)你可以提出一个强有力的、有根据的论点,即AI存在固有的缺陷或局限性,使其无法完成X,然后展示证据表明情况可能如此。4)你可以关注调节变量或中介变量:这种提示方式、方法、社会背景或联系影响了GPT-4完成X的能力,并且这需要成为未来关注的问题。5)你可以关注人类:人们对AI的反应如何?失败和成功是什么?它给我们带来了哪些危险、优势或变化?同样,这些并非详尽无遗,但总体而言,关于负面能力的声明远不如正面声明持久。

Ethan Mollick

@emollick

It is not inherently bad to publish research on the impact of AI that only refers to older models, but it requires a very careful discussion and has to be very clear to non-technical readers. If you show AI can do something, its fine. Generally, once AI has gained an ability, it does not regress in future generations. So papers that show that AI has crossed a threshold (such as "good as a human") or that seek to establish some sort of minimum impact or effect still hold up even if they used GPT-4. If the finding is that AI is bad or biased at something, you need to be much more careful. You cannot claim that because GPT-4 fails at something, that AI is bad at that task, because current or future models may do it. So what can you claim? These are just some examples: 1) You can just be explicit: "GPT-4 could not do X." This is a rare way to frame things because it is not really an interesting publishable paper in many cases. However, if you provide the information needed to reproduce your work, it can become a useful benchmark to measure progress. 2) You can measure trends: compare GPT-4 to GPT-5 to GPT-5.6 Sol or whatever (it is really important to include at least one reasoning model). You can then make better claims about abilities relating to this task. 3) You can make a strong, grounded argument that AI has a natural flaw or limitation that means it cannot do X, and then demonstrate evidence that this might be the case. 4) You can focus on a moderator or mediator: this sort of prompting or approach or social context or connection impacts the ability of GPT-4 to do X, and needs to be a concern for the future. 5) You can focus on humans: how do people react to AI? what are the failures and successes? what dangers or advantages or changes does it bring to us? Again, these aren't exhaustive, but generally negative capability claims have been much less durable than positive ones.

阅读原文