ProRL:延长强化学习扩展大语言模型的推理边界

ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models

刘明杰 Mingjie Liu · NVIDIA · 2025-05-30 · arXiv:2505.24864 ↗ · 被引 149

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

近期以推理为中心的语言模型进展凸显了强化学习(RL)作为将模型与可验证奖励对齐的一种有前景的方法。然而,关于 RL 是否真正扩展了模型的推理能力,还是仅仅放大了基础模型分布中已有的高奖励输出,以及持续扩大 RL 计算量是否可靠地提升推理性能,仍存在争议。在本工作中,我们通过展示延长 RL(ProRL)训练能够发现基础模型即使在大量采样下也无法触及的新推理策略,挑战了现有假设。我们提出了 ProRL,一种结合 KL 散度控制、参考策略重置和多样化任务套件的新型训练方法。我们的实证分析表明,RL 训练模型在广泛的 pass@k 评估中始终优于基础模型,包括基础模型无论尝试多少次都完全失败的情景。我们进一步表明,推理边界的改进与基础模型的任务能力和训练时长密切相关,这表明 RL 能够随时间探索和填充解空间的新区域。这些发现为 RL 在何种条件下有意义地扩展语言模型推理边界提供了新见解,并为未来关于长时域 RL 推理的研究奠定了基础。我们发布了模型权重以支持进一步研究:https://huggingface.co/nvidia/Nemotron-Research-Reasoning-Qwen-1.5B

Recent advances in reasoning-centric language models have highlighted reinforcement learning (RL) as a promising method for aligning models with verifiable rewards. However, it remains contentious whether RL truly expands a model's reasoning capabilities or merely amplifies high-reward outputs already latent in the base model's distribution, and whether continually scaling up RL compute reliably leads to improved reasoning performance. In this work, we challenge prevailing assumptions by demonstrating that prolonged RL (ProRL) training can uncover novel reasoning strategies that are inaccessible to base models, even under extensive sampling. We introduce ProRL, a novel training methodology that incorporates KL divergence control, reference policy resetting, and a diverse suite of tasks. Our empirical analysis reveals that RL-trained models consistently outperform base models across a wide range of pass@k evaluations, including scenarios where base models fail entirely regardless of the number of attempts. We further show that reasoning boundary improvements correlates strongly with task competence of base model and training duration, suggesting that RL can explore and populate new regions of solution space over time. These findings offer new insights into the conditions under which RL meaningfully expands reasoning boundaries in language models and establish a foundation for future work on long-horizon RL for reasoning. We release model weights to support further research: https://huggingface.co/nvidia/Nemotron-Research-Reasoning-Qwen-1.5B

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 19)

阅读逐段中英对照全文 →