DeepSeek-R1:通过强化学习激励大语言模型的推理能力

DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

郭达雅 Daya Guo · DeepSeek · 2025-01-22 · arXiv:2501.12948 ↗ · 被引 5417

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

通用推理是人工智能中一个长期且艰巨的挑战。最近的突破,例如大型语言模型(LLM)和思维链提示,在基础推理任务上取得了显著成功。然而,这种成功严重依赖于大量人工标注的示范,且模型的能力对于更复杂的问题仍显不足。本文表明,通过纯强化学习(RL)可以激励 LLM 的推理能力,从而无需人工标注的推理轨迹。所提出的 RL 框架促进了高级推理模式(如自我反思、验证和动态策略调整)的涌现发展。因此,训练后的模型在数学、编程竞赛和 STEM 领域等可验证任务上表现出色,超越了通过传统监督学习在人类示范上训练的同类模型。此外,这些大规模模型涌现出的推理模式可以被系统地利用,以指导和增强较小模型的推理能力。

General reasoning represents a long-standing and formidable challenge in artificial intelligence. Recent breakthroughs, exemplified by large language models (LLMs) and chain-of-thought prompting, have achieved considerable success on foundational reasoning tasks. However, this success is heavily contingent upon extensive human-annotated demonstrations, and models' capabilities are still insufficient for more complex problems. Here we show that the reasoning abilities of LLMs can be incentivized through pure reinforcement learning (RL), obviating the need for human-labeled reasoning trajectories. The proposed RL framework facilitates the emergent development of advanced reasoning patterns, such as self-reflection, verification, and dynamic strategy adaptation. Consequently, the trained model achieves superior performance on verifiable tasks such as mathematics, coding competitions, and STEM fields, surpassing its counterparts trained via conventional supervised learning on human demonstrations. Moreover, the emergent reasoning patterns exhibited by these large-scale models can be systematically harnessed to guide and enhance the reasoning capabilities of smaller models.

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 14)

阅读逐段中英对照全文 →