DeepSeek R1 复制 o1 的秘诀

DeepSeek R1's recipe to replicate o1

内森·兰伯特 Nathan Lambert · Allen Institute for AI · 2025-01-21 · Interconnects ↗

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

本文分析了 DeepSeek 发布的开源权重推理语言模型 R1 及其训练配方,该配方复现了 OpenAI 的 o1 方法。核心方法包括四个阶段:基于合成推理数据的冷启动监督微调、针对可验证问题的大规模强化学习、通过拒绝采样扩展通用能力,以及混合推理与偏好调优的最终强化学习阶段。作者强调,R1-Zero 作为仅使用强化学习、未经过监督微调的变体,是证明强化学习本身即可诱发推理行为的关键证据,尽管存在可用性问题。文章认为,此次开源发布标志着推理模型研究的转折点,从模糊的博客文章转向清晰、可复现的范式,并预测 2025 年将迎来快速进展和价格战。对读者而言,核心要点在于该配方的成功依赖于强大的基础模型、具有可验证奖励的强化学习,且技术创新并非护城河,因为像 R1 这样的开源模型能以极低的成本媲美专有模型。

This article analyzes DeepSeek's release of R1, an open-weights reasoning language model, and its training recipe, which replicates OpenAI's o1 approach. The core method involves a four-stage process: cold-start supervised finetuning on synthetic reasoning data, large-scale reinforcement learning (RL) on verifiable problems, rejection sampling to broaden general capabilities, and a final RL stage mixing reasoning and preference tuning. The author highlights R1-Zero, an RL-only variant trained without SFT, as a key proof that RL alone can induce reasoning behaviors, though with usability issues. The article argues that this open release marks a turning point in reasoning model research, shifting from opaque blog posts to a clear, reproducible paradigm, and predicts rapid progress and a price war in 2025. For readers, the key takeaway is that the recipe's success hinges on a strong base model, RL with verifiable rewards, and that technical innovations are not moats, as open models like R1 can match proprietary ones at a fraction of the cost.

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 8)

全文 · Full text(逐段中英对照)

是的,敲响真正的 o1 复制钟声给 DeepSeek R1 🔔🔔🔔。我们下一步走向何方。 Yes, ring the true o1 replication bells for DeepSeek R1 🔔🔔🔔. Where we go next.

_这周我有几个节目要与你分享:_

_I have a few shows to share with you this week:_

1. _一两周前在 The Retort 上,我们讨论了 AI 的本质以及它是否是一门科学(在库恩意义上)_

1. _On The Retort a week or two ago, we discussed the nature of AI and if it is a science (in the Kuhn’ian sense)_

2. _我出现在 Dean W. Ball 和 Timothy B. Lee 的新播客 AI Summer 上,讨论“思考模型”以及后训练与推理方法之间的边界。在此收听。_

2. _I appeared on Dean W. Ball_ _and Timothy B. Lee’s new podcast AI Summer to discuss “thinking models” and the border between post-training and reasoning methods. Listen here._

3. _最后,我在 NeurIPs 上关于我如何看待 AI 应用后训练的演讲现已公开。_

3. _Finally, a talk I gave at NeurIPs on how I think about post-training for AI applications is now public._

_这篇文章可能会在电子邮件收件箱中被截断——我建议点击标题在线阅读!_

_This post is likely getting cut off in email inboxes — I recommend reading online by clicking on the title!_

昨天,1 月 20 日,中国的开放权重前沿 AI 实验室 DeepSeek AI 发布了他们第一个完整的推理模型。它包含:

Yesterday, January 20th, China’s open-weights frontier AI laboratory, DeepSeek AI, released their first full fledged reasoning model. It came as:

* 一个旗舰推理语言模型 R1,通过一个 4 阶段、强化学习密集的过程训练。它采用 MIT 许可证,这意味着公司和研究人员可以基于其输出进行构建和训练,以加速推理语言模型(RLM)的开发和部署。

* A flagship reasoning language model, R1, trained via a 4-stage, RL heavy process. It is MIT-licensed which means companies and researchers can build upon and train on its outputs to accelerate the development and deployment of reasoning language models (RLMs).

* 一个仅使用强化学习的推理模型,直接从他们的 V3 基础模型训练而来,即 R1-Zero(用于为完整的 R1 创建训练数据)。

* An RL-only reasoning model trained directly from their V3 base model, R1-Zero (used to create training data for full R1).

* 一套使用从 R1 导出的监督微调(SFT)数据微调的开放权重模型(类似于他们中间训练步骤之一的数据)。

* A suite of open-weight models finetuned with supervised finetuning (SFT) data derived from R1 (similar data to one of their intermediate training steps).

* 一份详细说明其强化学习训练方法的技术报告。

* A technical report detailing their RL training methods.

* 模型可在 chat.deepseek.com(通过 DeepThink)和他们的新应用中使用。

* Models are available at chat.deepseek.com (via DeepThink) and in their new app.

这篇文章较少关注评估结果(当然,这些结果非常好,如下所示)1,而是更多关于_训练是如何进行的_以及_这一切意味着什么_。

This post is less about the evaluation results (which, of course, are extremely good and shown below)1, but rather about _how training is done_ and _what it all means_.

[](https://substackcdn.com/image/fetch/$s_!C1n1!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa2f93421-cd08-4cf9-b348-a5a20c7becaf_1476x1272.png)

[](https://substackcdn.com/image/fetch/$s_!C1n1!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa2f93421-cd08-4cf9-b348-a5a20c7becaf_1476x1272.png)

这是推理模型研究中不确定性的一个重大转折点。直到现在,推理模型一直是一个没有明确开创性论文的主要工业研究领域。在语言模型起飞之前,我们有像 GPT-2 论文用于预训练或 InstructGPT(以及 Anthropic 的白皮书)用于后训练。对于推理,我们只能盯着可能具有误导性的博客文章。推理研究和进展现在已被锁定——预计在 2025 年将取得巨大进展,并且更多进展将是开放的。

This is a major transition point in the uncertainty in reasoning model research. Until now, reasoning models have been a major area of industrial research without a clear seminal paper. Before language models took off, we had the likes of the GPT-2 paper for pretraining or InstructGPT (and Anthropic’s whitepapers) for post-training. For reasoning, we were staring at potentially misleading blog posts. Reasoning research and progress is now locked in — expect huge amounts of progress in 2025 and more of it in the open.

这再次证实了新的技术配方通常不是护城河——概念验证或泄露的动机通常会让知识传播出去。

This again confirms that new technical recipes normally aren’t moats — the motivation of a proof of concept or leaks normally get the knowledge out.

首先,看看这些推理模型的定价。OpenAI 可能因其长上下文服务的成本和作为市场上唯一模型而对其模型收取更高费用,但现在 o1 的定价为每百万输入 token 15 美元 / 每百万输出 token 60 美元,相对于 R1 的每百万输入 token 0.55 美元 / 每百万输出 token 2.19 美元显得格格不入(是的,o1-mini 更便宜,为每百万 token 3 美元/12 美元,但仍然有近 10 倍的差距)。即将到来的推理模型价格战将类似于 2023 年的 Mixtral 推理价格战。

For one, look at the pricing of these reasoning models. OpenAI was likely charging more for its model due to the costs of long-context serving and being the only model in town, but now o1’s pricing at $15 per million input tokens / $60 output looks out of place relative to R1’s pricing at $0.55 per million input tokens / $2.19 output (yes, o1-mini is cheaper at $3/$12 per million, but still almost a 10x difference). The price war that is coming for reasoning models will look like the Mixtral inference price war from 2023.

在 o3 方面,OpenAI 可能在技术上领先,但它并未普遍可用,而且权重也不会很快可用。这标志着自 Stable Diffusion 发布以来,最相关且被讨论最多的 AI 模型首次以非常友好的许可证发布。回顾过去 2.5 年“开源”AI 的历程,这是历史上一个令人惊讶的时刻。

With o3, OpenAI is likely technically ahead, but it is not generally available nor will the weights be available anytime soon. This points to the first time since Stable Diffusion’s release that the most relevant and discussed AI model is released with a very friendly license. Looking back at the journey “open-source” AI has been on over the last 2.5 years, this is a surprising moment in time marked in the history books.

我们不完全知道这些模型未来将如何用于代码和数学之外,但不断有声音表明 OpenAI 的 o1-Pro 是许多更具挑战性任务的最佳模型(我需要亲自尝试后才能做出明确推荐)。

We don’t entirely know how these models will be used in the future beyond code and math, but noises are constantly bubbling up that OpenAI’s o1-Pro is the best model for many more challenging tasks (I need to try it myself before making definitive recommendations).

现在最有用的文章是建立研究领域、注意事项和开放问题的文章。让我们深入细节。

The most useful post to write now is one that establishes the research area, the do’s and don’ts, and the open questions. Let’s get into the details.

DeepSeek R1 推理训练方案 The DeepSeek R1 training recipe for reasoning

1. 在来自 R1-Zero 模型的合成推理数据上进行监督微调的“冷启动”。

1. “Cold-start” of supervised finetuning on synthetic reasoning data from the R1-Zero model.2

2. 在推理问题上进行大规模强化学习训练,“直至收敛”。

2. Large-scale reinforcement learning training on reasoning problems “until convergence.”

3. 对 3/4 的推理问题和 1/4 的通用查询进行拒绝采样,以开始向通用模型的过渡。

3. Rejection sampling on 3/4 reasoning problems and 1/4 general queries to start the transition to a general-purpose model.

4. 混合推理问题(可验证奖励)与通用偏好调整奖励模型进行强化学习训练,以打磨模型。

4. Reinforcement learning training mixing reasoning problems (verifiable rewards) with general preference tuning reward models to polish the model.

下文将详细分解每个训练阶段的核心组成部分、见解和未解决问题。

Below, the post breaks down each training stage into its core components, insights, and open questions.

o1 复现的风向已强烈偏离任何形式的显式搜索(尤其是在推理时)。它过去是、现在也仍然是一个语言模型,其新的推理行为来自大量的强化学习训练。

The winds of o1 replication have been blowing strongly away from any sort explicit search (especially at inference time). It really was, and is, a language model with the new reasoning behaviors coming from a lot of RL training.

#### OpenAI 的 o1 使用“搜索”是一场心理战 [Nathan Lambert · 2024 年 12 月 4 日 阅读全文](https://www.interconnects.ai/p/openais-o1-using-search-was-a-psyop)

#### OpenAI's o1 using "search" was a PSYOP [Nathan Lambert · December 4, 2024 Read full story](https://www.interconnects.ai/p/openais-o1-using-search-was-a-psyop)

在开始之前,请记住,要做好这种推理训练,你需要一个具有长上下文能力的非常强大的基础模型。与标准的后训练类似,我们并不真正了解基础模型的哪些特质使其更适合直接进行强化学习训练。

Before we start, remember that to do this reasoning training well you need a very strong base model with long-context capabilities. Much like for standard post-training, we don’t really know what traits of a base model make for one that is more suited for direct RL training.

步骤 0:训练 R1-Zero,用合成数据初始化 R1 Step 0. Training R1-Zero to initialize R1 with synthetic data

DeepSeek R1 Zero 将作为首个在没有监督微调(SFT)作为初步步骤的情况下,通过“大规模强化学习(RL)”训练的开源模型而闻名。此前有传言提到 o1 也采用了类似方法,但其工作原理尚不明确。这是一个奇特的模型,DeepSeek 报告称它有时会在推理过程中切换语言,或表现出其他可靠性问题。

DeepSeek R1 Zero will be best known as the first open model trained with “large-scale reinforcement learning (RL) without supervised fine-tuning (SFT) as a preliminary step.” Rumors had mentioned this for o1, but understanding how it worked wasn’t clear. This is a funky model that DeepSeek reports will sometimes change languages in reasoning or show signs of other reliability issues.

R1-Zero 的轻微可用性问题表明,训练一个出色的推理模型不仅需要大规模强化学习,但 RL 部分确实是解锁我们寻求的推理行为的关键。

The minor usability issues with R1-Zero show why more than _just_ large-scale RL is needed to train a fantastic reasoning model, but the RL part is the key to unlocking the reasoning behaviors we are searching for.

他们展示了 R1-Zero 最有趣的结果,包括我一直期待的 RL 训练时间缩放图。自 o1 发布以来,每个人都痴迷于展示推理时间与评估性能相关性的图表。推理时间更容易通过蒙特卡洛树搜索等框架来诱发(或强制实现),但通过 RL 展示训练时间的改进才是真正的基础性成果。这正是我在研究中寻找的结果。

They include the most interesting results for R1-Zero, including the plot I’ve been asking for of RL-training time scaling. Since o1’s release, everyone has been obsessed with the plots showing how _inference time_ is correlated with evaluation performance. Inference time is far easier to elicit (or force by using a framework like Monte Carlo Tree Search), but showing _training time_ improvements via RL is the real foundational result. This is the result I’m searching for in my research.

[](https://substackcdn.com/image/fetch/$s_!8mIS!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7cc1c572-f0ab-4247-a062-04780a3321ed_1916x1074.png)

[](https://substackcdn.com/image/fetch/$s_!8mIS!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7cc1c572-f0ab-4247-a062-04780a3321ed_1916x1074.png)

还有一个意料之中但令人满意的图表,展示了生成长度随训练增长。这可以与上面的图表结合,形成我们见过的许多版本但方法不那么清晰的“推理时间缩放”图之一。

And an unsurprising, yet very satisfying plot of length growing with training. This could be mixed with the above plot to make one of the “inference time scaling” plots we have seen many versions of with less clear methods.

[](https://substackcdn.com/image/fetch/$s_!KVEL!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F75974ec6-d89b-4c81-a7e5-9530145f5e6c_1910x1088.png)

[](https://substackcdn.com/image/fetch/$s_!KVEL!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F75974ec6-d89b-4c81-a7e5-9530145f5e6c_1910x1088.png)

在这两个图中,如果让 RL 继续训练更长时间,数值似乎还能继续上升。由于进展速度如此之快,这些实验室通过让训练在接近饱和时结束并开始下一个实验来获得更多收益,而不是追求那最后的 1%。

In both of these plots, it looks like the numbers could still be going up if they let the RL cook longer. With the pace of progress so high, these laboratories get more gains by ending the jobs near saturation and starting the next experiment instead of seeking that last 1%.

大多数(如果不是全部)研究人员会跳过训练 R1-Zero 风格模型的步骤,因为他们不需要。DeepSeek 明确表示,他们的 SFT 推理轨迹“冷启动”使最终的 R1 模型更好——这并不意外,因为他们希望 R1 成为某种特定类型的指令微调模型。这有助于避免 R1-Zero 中 DeepSeek 提到的一些“RL 异常”,例如在生成过程中切换语言。

Most, if not all, researchers will skip the step of training an R1-Zero style model because they don’t need to. DeepSeek made it clear that their “cold start” of SFT reasoning traces makes the final R1 model better — this is unsurprising, as they want R1 to be a certain type of instruction-tuned model. It’ll help avoid some of the “RL oddities” in R1-Zero that DeepSeek mentions like changing language mid-generation.

尽管如此,基于基础模型的 RL 领域应进一步研究。R1-Zero 的训练方式相当巧妙,因为大多数没有经过任何指令微调的基础模型都存在严重问题,比如胡言乱语且从不生成停止标记。R1-Zero 通过系统提示让模型生成 <answer> HTML 标签来避免这个问题。此外,我怀疑这种训练方式在较旧的基础模型上无法奏效,因为这些模型的预训练语料中没有包含一些标准的后训练风格指令数据。例如,在 OLMo 2 中,我们在退火混合数据中加入了一些 MATH 指令数据。只需少量指令就能让这个系统提示生效。

Still, the area of RL-on-base-models should be studied further. The way that R1-Zero can be trained is quite clever as most base models without any instruction tuning have a major issues with rambling and never generating a stop token. R1-Zero avoids this with a system prompt telling the model to generate <answer> HTML tags. Additionally, I suspect this type of training wouldn’t work on older base models that don’t have some standard post-training style instruction data in the pretraining corpus. For example, in OLMo 2 we had some MATH instruction data in the annealing mix. Just a few instructions will let this system prompt work.

事实上,当直接从基础模型训练而非从标准后训练模型(没有冗长的思维链风格)训练时,通过 RL 训练增加生成长度的趋势可能更强。为了让 RL 在这样的指令遵循模型中真正开始大幅增加响应长度,它必须摆脱已固化的特定响应长度。例如,在 Tülu 3 的 RL 微调最后阶段,响应率首先下降的阶段可能是较大规模的 SFT 训练与较小规模的 RL 设置之间不匹配的障碍。

In fact, the trend of increasing generation length via RL training could be even stronger when training directly from a base model rather than a standard post-trained model that doesn’t have a verbose chain of thought style. In order for RL to really start cranking up the response length in such an instruction-following model it will have to unlearn a certain response length that was baked in. For example, in Tülu 3’s final stage of RL finetuning, the phase where the response rate first goes down could be the barrier of misalignment between a larger round of SFT training before a smaller RL setup.

[](https://substackcdn.com/image/fetch/$s_!2vF9!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8dab8b2e-ee61-41d6-911a-bcc105d7f5da_1564x356.png)

[](https://substackcdn.com/image/fetch/$s_!2vF9!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8dab8b2e-ee61-41d6-911a-bcc105d7f5da_1564x356.png)

注意,这里的 x 轴是 episodes,与 DeepSeek 图中使用的“step”不同。

Note, the x axis here is episodes, which is different than the “step” used in the DeepSeek plots.

放大这些 R1-Zero 图的 x 轴,你可以看到它们进行了数千个“RL 步骤”。这里的 RL 步骤指的是模型更新步骤,该步骤在批次中为提示生成多个输出并验证答案之后进行。3 这是大量的 RL 训练,尤其是对于如此大的模型。作为参考,在我们的 Tülu 3 工作中,我们通常对模型进行数百步的微调,而我们即将发布的最大模型仅训练了约 50 步的 RL。

Zooming in on the x-axes of these R1-Zero plots, you can see that they’re doing 1000s of “RL steps.” RL step in this case refers to the model update step, which comes after multiple generations are made for the prompts in the batch and then answers are verified.3 This is a large amount of RL training, especially with such a large model. For reference, in our Tülu 3 work, we finetuned our models for 100s of steps normally, and the biggest models we are releasing soon only trained for ~50 steps of RL.

相对于现有文献,这是规模扩大的 RL。R1 本身肯定使用了类似的设置,但 DeepSeek 没有包含相同的细节,因此本文其余部分更多地依赖论文中的明确文本。

This is scaled-up RL relative to existing literature. R1 proper surely uses a similar setup, but DeepSeek did not include the same details, so the rest of this post relies more on explicit text in the paper.

步骤 1:推理 SFT“冷启动” Step 1. Reasoning SFT “Cold Start”

为了提高可读性(即帮助保持格式)并提升最终推理模型的性能,DeepSeek 在原始基础模型上使用来自 R1-Zero 模型的“数千个”筛选完成结果进行少量监督微调。这涉及一些技巧(似乎没有哪个是必不可少的,你只需要一些这样的数据),例如:

In order to improve the readability (i.e. help maintain formatting) and increase the final performance of the final reasoning model, DeepSeek performs a small amount of supervised finetuning on the original base model with “a few thousand” filtered completions from the R1-Zero model. This involves a few tricks (none of which seem essential, you just need some of this data), such as:

对于复现工作,这些方法中的任何一种都可以采用。事实上,使用 DeepSeek-R1 本身可能是最简单的方法。

For replication efforts, any of these can be done. In fact, using DeepSeek-R1 itself is likely the easiest way.

这一阶段为模型的损失景观做好准备,使得在强化学习训练中更容易出现诸如“等等,让我检查一下我的工作”或“那是错误的”这样的“涌现”行为。

This phase readies the loss landscape of the model to make the “emergent” behaviors like “wait, let me check my work” or “that was wrong” come forth more easily in RL training.

步骤 2. 用于推理的大规模强化学习 Step 2. Large-scale RL for reasoning

提醒一下,用于推理模型的强化学习建立在一个简单的想法上:对于可以检查正确答案的问题,当模型给出正确答案时给予奖励。其基本反馈循环如下所示:

As a reminder, RL for reasoning models is built on a simple idea where you should reward the model for getting correct answers to problems where you can check if it has a correct answer. A basic feedback loop of this looks like the following:

[](https://substackcdn.com/image/fetch/$s_!4T2d!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff1b709d1-9bb7-436f-9d07-08ead28eef2d_923x477.png)

[](https://substackcdn.com/image/fetch/$s_!4T2d!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff1b709d1-9bb7-436f-9d07-08ead28eef2d_923x477.png)

Tülu 3 中 RLVR 微调的系统图。https://arxiv.org/abs/2411.15124

System diagram for the RLVR finetuning in Tülu 3. https://arxiv.org/abs/2411.15124

这里的“奖励”具体是什么(同样的问题也适用于 R1-Zero)并未详细说明。DeepSeek 在强化学习的推理阶段提到了三个奖励组成部分:

Exactly what the “reward” is here (the same question applies for R1-Zero) isn’t detailed. DeepSeek mentions three reward components during the reasoning phase of RL:4

1. 准确性奖励:如果对提示的响应正确,则给予分数奖励。我将其称为“可验证”领域,在 OpenAI 的强化微调中,这由其评分器处理。简而言之:如果答案正确,奖励为正;否则为 0。

1. Accuracy rewards: These are score bonuses if the response to a prompt is correct. I’ve been referring to these as “verifiable” domains and in OpenAI’s Reinforcement Finetuning this is handled by their graders. TLDR: If the answer is correct, the reward is positive, if not, it is 0.5

2. 格式奖励:这些奖励(如果不满足则为惩罚)用于检查并确保模型遵循正确的格式,如 <think> 和 </think> 以及 <answer> 和 </answer>,以实现稳定的推理。

2. Format rewards: These are rewards (or penalties if not satisfied) to check and make sure that the model follows the correct formatting of <think> or </think> and <answer> and </answer> for stable inference.

3. 语言一致性奖励:如果答案的语言与问题的语言 100% 匹配,则给模型添加奖励。DeepSeek 写道,这一额外奖励虽然导致模型性能“略有下降”,但更符合人类偏好。添加它是为了让模型更易于使用,这很好地提醒我们评估分数并非一切。

3. Language consistency rewards: A reward is added to the model if the language of the answer is 100% matching the language of the question. DeepSeek writes that this additional reward shows a “slight degradation in the model’s performance,” but better human preferences. It’s added to make the model nice to use, which is a wonderful reminder that evaluation scores are not all that matters.

这里的第一个奖励驱动了大部分学习,而另外两个则是创建稳定模型的护栏(这并不是说它们不是重要的实现细节,而是说第一个是必要的,而其他可能不是)。为了优化这个奖励,DeepSeek 使用了他们引入的强化学习算法——组相对策略优化(GRPO),这是一种 PPO 更新规则,但使用基于蒙特卡洛优势估计的不同值近似方法,而不是在内存中维护单独的值模型。选择这一方法最可能的解释(就像 OpenAI 一直使用 PPO 一样)是它在他们的基础设施中已经成熟实现。

The first reward here drives the majority of the learning and the other two are guardrails for creating a stable model (which is not to say they aren’t important implementation details, but rather that the first one is necessary and the others may not be). To optimize this reward, DeepSeek uses the RL algorithm that they introduced, Group Relative Policy Optimization, which is the PPO update rule with a different value approximation method based on Monte Carlo advantage estimates rather than holding a separate value model in memory. The most likely explanation for this choice (much like how OpenAI has always used PPO) is that it is the mature implementation in their infrastructure.

这张来自 DeepSeekMath 论文的图片是 PPO 与 GRPO 的绝佳比较(如果你只关心整体流程,可以跳过此部分):

This image from the DeepSeekMath paper is a fantastic comparison of PPO to GRPO (this is fine to skip this if you only care about the big picture recipe):

[](https://substackcdn.com/image/fetch/$s_!douJ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F21ce968f-4017-498f-a286-1f0ef034fa2a_883x441.png)

[](https://substackcdn.com/image/fetch/$s_!douJ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F21ce968f-4017-498f-a286-1f0ef034fa2a_883x441.png)

奖励设置(以及数据)的性质是这种推理训练的关键,许多小的强化学习细节可以相互替代。

The nature of the reward setup (and the data) is the key to this sort of reasoning training and many of the small RL details can be substituted for each other.

与 DeepSeek V3 论文类似,这里没有包含他们用于训练模型的数据细节。这绝对至关重要,并且几乎肯定涉及许多许多带有答案的可验证提示。为了研究这些模型,社区需要这些数据集的开放版本。

Much like the DeepSeek V3 paper, the details of what data they used to train the model are not included here. This is absolutely crucial and almost certainly involves many, many verifiable prompts with answers. In order to study these models the community needs open versions of these datasets.

我很想看到他们强化学习基础设施的细节(类似于 DeepSeek V3 论文中的细节),因为许多人正试图在这些模型上进行构建。强化学习训练需要在内存中维护多个模型,并在生成、验证和计算损失步骤之间交替。正如 Sasha Rush 所说,“我们需要尽快编写验证器”,这正是我们在 Ai2 基于 Tülu 3 试图做的事情,并且非常需要开源代码方面的帮助。对于感兴趣的相关方,一个好的方法是逐个领域开发工具和数据。

I would’ve loved to see details of their RL infrastructure (similar to the details in the DeepSeek V3 paper), as many people are looking to build on these models. RL training requires holding multiple models in memory and alternating between generating, verifying, and taking loss steps. As Sasha Rush says, “We need to code up verifiers ASAP,” which is what we are trying to do at Ai2 building on Tülu 3 and could use a lot of help with the open-source code.6 A good approach for entities interested here is to develop tooling and data for one domain at a time.

前两个步骤并不新鲜,而是人们广泛讨论过的想法的规模化版本。DeepSeek 在论文中详述的最后两个步骤是已知技术的新应用,旨在利用其原始推理性能并“训练一个用户友好的模型”。

These first two steps are not new but rather scaled-up versions of ideas people have been discussing extensively. The final two steps DeepSeek details in the paper are new applications of known techniques to help take their raw reasoning performance and “train a user-friendly model.”

步骤 3. 拒绝采样以引入通用能力 Step 3. Rejection Sampling to introduce general abilities

拒绝采样是一种技术,你从模型生成补全,通过奖励模型对其进行排序,然后微调原始模型(通常使用监督微调损失)以提升在多种任务上的性能。这是 Llama 3 和许多其他模型使用的标准后训练工具之一。

Rejection sampling is a technique where you generate completions from a model, rank them via a reward model, and then finetune the original model (normally with the supervised finetuning loss) to improve performance on a variety of tasks. It’s one of the standard post-training tools used by Llama 3 and many others.

DeepSeek 使用拒绝采样开始将通用能力重新引入模型。这也是唯一一个包含数据数量的阶段——总共 80 万次补全,分为 60 万用于推理和 20 万用于通用聊天问题。考虑到这只是后期阶段的 SFT 训练,80 万这个数字对我来说并不意外,但它与我们在 Tülu 3 SFT 混合中使用的约 100 万提示的大小相似,这是领先后训练配方的数量级。

DeepSeek uses rejection sampling to begin to introduce general capabilities back into the model. It is also the one stage where they include data numbers — 800K completions total, split as 600K for reasoning and 200K for general chat problems. The 800K number is not surprising to me given this is just a late-stage SFT training, but it is similar in size to the ~1M prompts we used in the Tülu 3 SFT mix which is the ballpark for leading post-training recipes.

论文中的细节主要围绕为提示生成响应和过滤以优先选择高质量训练数据的方法。为了将更多领域纳入模型的能力范围,DeepSeek 采用了多种技巧,例如:

The details in the paper are largely around methods for _generating responses_ to prompts and _filtering_ to prioritize high-quality training data. In order to bring more domains into the scope of abilities for the model, DeepSeek has a variety of tricks, such as:

* 使用生成式奖励模型(即 LLM 作为评判者)来验证可能无法明确验证的问题的答案,

* Using generative reward models (i.e. LLM-as-a-judge) to verify answers to questions that may not be explicitly verifiable,

* 来自 DeepSeek-V3 标准后训练流程的数据,以及

* Data from the DeepSeek-V3 standard post-training pipeline, and

* 标准(不可验证)聊天数据,在回答前通过扩展思维链进行增强,以帮助模型将推理训练泛化到更广泛的用例。

* Standard (nonverifiable) chat data augmented with extended chain of thought before answering to help the model generalize from reasoning training to broader use cases.

总而言之,我们目前对此的细节知之甚少,并且存在很大的学习(以及可能的改进)空间。

All in, we currently have very few details here and there is a lot of open space to learn (and likely improve).

第 4 步:面向通用用途的最终强化学习训练 Step 4. Final RL training for general use

最后,DeepSeek R1 回归强化学习,这似乎是当前大多数微调工作的终点。第二阶段 RL 的目标是“在提升模型有用性和无害性的同时,进一步优化其推理能力”。

Finally, DeepSeek R1 goes back to reinforcement learning, which really seems to be how most finetuning is ending these days. The second RL stage is “aimed at improving the model’s helpfulness and harmlessness while simultaneously refining its reasoning capabilities.”

为此,他们进行了混合提示的 RL 训练:一部分来自可验证领域(如 R1-Zero 的做法),另一部分用于标准的基于人类反馈的强化学习(RLHF)偏好调优。他们使用了多个奖励模型,并基于 DeepSeek V3 的后训练配方进行构建。

In order to do this, they do RL training that mixes prompts from the verifiable domains (as done for R1-Zero) and prompts for standard RLHF preference tuning. In order to do this they have multiple reward models and build upon their post-training recipe in DeepSeek V3.

这并不容易实现,涉及许多问题:数据平衡如何把握?能否直接使用现成的奖励模型,还是需要它见过长推理轨迹?是否需要额外步骤来防止性能下降?等等。

This is not easy to do and involves many questions: What is the right data balance? Can you use an off-the-shelf existing reward model or does it need to have seen long reasoning traces? Are there additional steps needed to not degrade performance? And so on.

随着这一领域的研究和发展不断深入,这些问题将逐步得到解答。

As this grows into a larger area of research and development these questions will slowly be answered.

随着本文进入训练后期阶段,显然许多细节尚不明确。我们掌握了流程的大致框架,后续将逐步填充细节。我手头有一长串与推理相关的研究论文需要研读,尽管它们发表于 DeepSeek R1 之前,但仍能指向答案。

As this post has transitioned into the later stages of training, it is clear that many details are unknown. We have the general shape of how to sequence things and will fill in the details from here. I have a very long stack of reasoning-related research papers to poke through, and while they came before DeepSeek R1, they still will point toward answers.

所有这些问题都是可解的,DeepSeek 从 o1 发布到以开放权重模型达到同等性能的速度就是证明。

All of this is solvable, as proven by how quickly DeepSeek went from the o1 release to matching performance with an open weights model.

Interconnects 是由读者支持的出版物。欢迎考虑订阅。

Interconnects is a reader-supported publication. Consider becoming a subscriber.

讨论与下一步 Discussions and next steps

DeepSeek R1 报告有一个完整的子节专门讨论其蒸馏实验,即从 R1 模型获取输出,并用这些输出微调现有的开源权重模型以提升性能。他们发布这些内容是一项极好的服务,并为较小模型上的强化学习实验提供了坚实的基线,以便在不久的将来尝试匹配。

The DeepSeek R1 report has an entire other subsection dedicated to its distillation experiments, where it took completions from the R1 model and finetuned existing open-weight models with them to boost performance. This is a fantastic service for them to release this and provides a solid baseline for RL experiments on smaller models to try and match in the near future.

论文中关于需要大型模型才能获得最大的推理收益(并生成有效的合成数据)的讨论,可能是最大的开放性问题:

The discussion in the paper on how large models are required to see the biggest reasoning gains (and generate effective synthetic data) is likely the biggest open question:

随着小型模型逐年改进,同样的训练方式很可能适用于像 Llama 5 或 6 8B 这样的模型。这给我们留下了一个同样开放的问题:为什么不同的能力会在更大的模型上“涌现”?缩放定律是每一代前沿模型往往都是可用最大模型的原因。2025 年这个问题的激动人心之处在于:语言建模研究的缓慢进展会将高级推理能力驱动到多小的模型上?

As smaller models continually improve over the years, it is likely that the same type of training could work on something like Llama 5 or 6 8B. It leaves us with the same, open question as to _why_ different abilities “emerge” at larger models. Scaling laws are the reasons that each generation’s _frontier_ models tend to be the largest models available. The exciting form of this question for 2025 is: How _small_ will the slow progress of language modeling research drive advanced reasoning capabilities?

每隔一段时间,就会有一篇论文让前进的道路变得清晰。上一次我有这种感觉是 Llama 3 报告关于后训练的部分,后来被整合到 Tülu 3 论文中。

Every so often a paper comes around that makes the path forward clear. The last time I felt this way was with the Llama 3 report for post-training, which solidified into the Tülu 3 paper.

#### 前沿模型后训练的配方 [Nathan Lambert · 2024 年 8 月 7 日 阅读全文](https://www.interconnects.ai/p/frontier-model-post-training)

#### A recipe for frontier model post-training [Nathan Lambert · August 7, 2024 Read full story](https://www.interconnects.ai/p/frontier-model-post-training)

* 推理轨迹的蒸馏(如 R1 论文中所做),

* Distillation of reasoning traces (as done in the R1 paper),

* 过程奖励模型(PRM)和蒙特卡洛树搜索(MCTS)的消亡,

* The demise of process reward models (PRMs) and Monte Carlo Tree Search (MCTS),

* DeepSeek 论文中一些让我恼火的内容,比如“啊哈”时刻和过度依赖人类先验,

* Some things in the DeepSeek paper, like the “Aha” moment and over-indexing on human priors, that annoy me,

* 来自学术界的新推理研究,

* The new reasoning research coming out from academia,

* 昨天发布的另一个推理模型——Kimi 1.5,

* The other reasoning model that dropped yesterday — Kimi 1.5,

* Tülu 3 RLVR 迄今为止最大的应用,以及

* The biggest application of Tülu 3 RLVR yet, and

* 推理模型领域正在争论的所有其他想法。

* All the other ideas that are under debate in the reasoning model space.

R1 肯定不是训练这些模型的唯一方法,但它是人们将立即在此基础上构建的配方。让我们开始着手更多的数据集和基础设施。

R1 is surely not the only way to train these models, but it is the recipe that people will build off immediately. Let’s get cranking on more datasets and infrastructure.

_新读者可以查看 Interconnects 上的推理与推理标签!_

_For those new here, you can check out the Inference & Reasoning tag on Interconnects!_

要确信模型实际有多接近,需要大量的工作。如果你非要我选,我会选 OpenAI,因为他们服务的用户群更大,而且在我看来,他们在结构上不太倾向于最大化评估分数而忽视实际有用性。随着时间的推移,我会在这里形成更强烈的观点,但初步的对话让我觉得 R1 模型已经接近了。

It will take a substantial amount of work to be confident in how close the models actually are. If you had to have me choose I would choose OpenAI, as they’re serving a larger user base and in my opinion are less structurally inclined to maximize evaluation scores relative to actual usefulness. I’ll formulate stronger opinions here over time, but initial conversations make me think that the R1 model is in the ballpark.

我不把生成推理数据算作第 0 步,因为现在可以轻松地用 R1、Qwen QwQ 或替代方案完成。很快,HuggingFace 上就会有大量数据。

I do not count generating the reasoning data as step 0 as this can now be done easily with R1, Qwen QwQ, or alternatives. Soon, there will be plenty on HuggingFace.

在我们的 Tülu 3 论文中,我们通常报告 episodes,这与 steps 相关,因为它是我们生成并验证的总提示数。episodes 的数量与批量大小乘以强化学习步数成正比。

In our Tülu 3 paper we generally have reported episodes, which is related to steps, as it is the total numbers of prompts that we generate for and verify. The number of episodes is proportional to the batch size times the number of RL steps.

DeepSeek 说他们“直接求和”这些奖励,但不知道尺度和形状,这并不特别有启发性。

DeepSeek says they “directly sum” these rewards, but without knowing the scales and shapes, that is not particularly insightful.

Ross Taylor 也在询问他们的奖励塑造,这与验证器密切相关。

Ross Taylor is also asking about their reward shaping, which interacts very closely with the verifiers.

互动版:图/公式 + 针对本篇提问 →