arXiv cs.LG论文
测量与缓解 RLVR 中的解模式坍缩
该论文指出,在可验证奖励强化学习(RLVR)中,模型对正确答案的具体形式不敏感,导致训练过程中出现“解模式坍缩”,即模型倾向于重复已知答案而丧失多样性。作者提出通过测量和缓解这一现象,使模型在训练中保留多种正确的解决方案,从而提升模型的泛化能力和鲁棒性。
这篇是正式发表的长论文,站内提供中文解读,全文请到原文阅读 PDF。
Abstract:A language model (LM) can usually answer the same question in more than one way, but reinforcement learning with verifiable rewards (RLVR) is indifferent to which correct answer a model produces. A solution will earn the same reward whether it is the thousandth copy of a familiar answer or one the model has never produced before. Yet, there is potential value in having the model retain multiple correct solutions as it is trained. For instance, multiple modes may give users a choice and provide problem-solving strategies that improve overall model performance. Here, we introduce ModeBench, a benchmark of multi-solution tasks in which the verifier returns both correctness and mode discovered. We then use ModeBench to measure how solution diversity changes under RLVR post-training. We find that RLVR post-training concentrates probability onto fewer correct modes even as accuracy holds or improves, and moreover, that frontier models are already highly concentrated. We then introduce our solution, Re:Max, which stores one verified example per discovered mode in a replay buffer and trains on those stored modes uniformly. A solution found once is, therefore, practiced as often as one found repeatedly. Across three model scales, two RL objectives, and harder task constructions, replay improves both how often a policy succeeds and how many different ways it can succeed.