Understanding Reasoning LLMs
打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→本文全面概述了大语言模型(LLM)领域中的推理模型,将其定义为擅长处理谜题、高等数学和编程挑战等复杂多步任务的系统。文章概述了构建和增强推理能力的四种主要方法:推理时扩展、纯强化学习、监督微调结合强化学习以及蒸馏。文章以 DeepSeek 的 R1 流水线为案例研究,详细介绍了其变体(R1-Zero、R1 和 R1-Distill)的开发过程,强调了在没有初始监督微调的情况下,纯强化学习意外涌现出推理能力。文章还讨论了推理模型的优势与局限,指出其成本较高且输出冗长,并建议仅在任务确实需要复杂推理时使用。结论强调,虽然推理模型是一种有价值的专门化工具,但并非万能解决方案,模型的选择应与任务的复杂性相匹配。
This article provides a comprehensive overview of reasoning models in the field of large language models (LLMs), defining them as systems that excel at complex, multi-step tasks such as puzzles, advanced mathematics, and coding challenges. It outlines four primary approaches to building and enhancing reasoning capabilities: inference-time scaling, pure reinforcement learning, supervised fine-tuning combined with RL, and distillation. The article uses DeepSeek's R1 pipeline as a case study, detailing how its variants (R1-Zero, R1, and R1-Distill) were developed, highlighting the surprising emergence of reasoning from pure RL without initial SFT. It also discusses the strengths and limitations of reasoning models, noting their higher cost and verbosity, and advises using them only for tasks that genuinely require complex reasoning. The conclusion emphasizes that while reasoning models are a valuable specialization, they are not a universal solution, and the choice of model should align with the task's complexity.
本文介绍了构建推理模型的四种主要方法,即如何增强大语言模型(LLM)的推理能力。希望这能提供有价值的见解,并帮助你在快速发展的文献和围绕该话题的热议中导航。
This article describes the four main approaches to building reasoning models, or how we can enhance LLMs with reasoning capabilities. I hope this provides valuable insights and helps you navigate the rapidly evolving literature and hype surrounding this topic.
2024 年,LLM 领域出现了日益明显的专业化趋势。除了预训练和微调,我们见证了从 RAG 到代码助手等专业化应用的兴起。我预计这一趋势将在 2025 年加速,更加注重领域和应用的特定优化(即“专业化”)。
In 2024, the LLM field saw increasing specialization. Beyond pre-training and fine-tuning, we witnessed the rise of specialized applications, from RAGs to code assistants. I expect this trend to accelerate in 2025, with an even greater emphasis on domain- and application-specific optimizations (i.e., "specializations").
[](https://substackcdn.com/image/fetch/$s_!QwUc!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd6ebc5c9-461f-4d3a-889b-b8ea4e14e5ba_1600x830.png)
[](https://substackcdn.com/image/fetch/$s_!QwUc!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd6ebc5c9-461f-4d3a-889b-b8ea4e14e5ba_1600x830.png)
_图 1:阶段 1-3 是开发 LLM 的常见步骤。阶段 4 针对特定用例对 LLM 进行专业化。_
_Figure 1: Stages 1-3 are the common steps to developing LLMs. Stage 4 specializes LLMs for specific use cases._
推理模型的开发就是这些专业化方向之一。这意味着我们优化 LLM,使其在需要中间步骤的复杂任务(如谜题、高等数学和编程挑战)中表现出色。然而,这种专业化并不会取代其他 LLM 应用,因为将 LLM 转变为推理模型也会带来某些缺点,我将在后面讨论。
The development of reasoning models is one of these specializations. This means we refine LLMs to excel at complex tasks that are best solved with intermediate steps, such as puzzles, advanced math, and coding challenges. However, this specialization does not replace other LLM applications. Because transforming an LLM into a reasoning model also introduces certain drawbacks, which I will discuss later.
为了让您对下文内容有一个简要的了解,在本文中,我将:
To give you a brief glimpse of what's covered below, in this article, I will:
1. 解释“推理模型”的含义
1. Explain the meaning of "reasoning model"
2. 讨论推理模型的优缺点
2. Discuss the advantages and disadvantages of reasoning models
3. 概述 DeepSeek R1 背后的方法论
3. Outline the methodology behind DeepSeek R1
4. 描述构建和改进推理模型的四种主要方法
4. Describe the four main approaches to building and improving reasoning models
5. 分享对 DeepSeek V3 和 R1 发布后 LLM 格局的看法。
5. Share thoughts on the LLM landscape following the DeepSeek V3 and R1 releases.
6. 提供在预算紧张的情况下开发推理模型的技巧。
6. Provide tips for developing reasoning models on a tight budget.
随着今年 AI 的快速发展,希望这篇文章对您有所帮助!
I hope you find this article useful as AI continues its rapid development this year!
如果你从事人工智能(或一般机器学习)工作,你可能熟悉那些模糊且争议不断的定义。“推理模型”一词也不例外。最终,会有人在论文中正式定义它,但紧接着又会被重新定义,如此循环往复。
If you work in AI (or machine learning in general), you are probably familiar with vague and hotly debated definitions. The term "reasoning models" is no exception. Eventually, someone will define it formally in a paper, only for it to be redefined in the next, and so on.
在本文中,我将“推理”定义为回答需要复杂、多步骤生成并包含中间步骤的问题的过程。例如,事实性问答如“法国的首都是什么?”不涉及推理。相反,像“如果一列火车以 60 英里/小时的速度行驶 3 小时,它走了多远?”这样的问题则需要一些简单的推理。例如,它需要先认识到距离、速度和速度之间的关系,然后才能得出答案。
In this article, I define "reasoning" as the process of answering questions that require complex, multi-step generation with intermediate steps. For example, factual question-answering like "What is the capital of France?" does not involve reasoning. In contrast, a question like "If a train is moving at 60 mph and travels for 3 hours, how far does it go?" requires some simple reasoning. For instance, it requires recognizing the relationship between distance, speed, and time before arriving at the answer.
图 2:常规 LLM 可能只提供简短答案(如左图所示),而推理模型通常包含中间步骤,揭示部分思考过程。(注意,许多未专门针对推理任务开发的 LLM 也能在其答案中提供中间推理步骤。)
Figure 2: A regular LLM may only provide a short answer (as shown on the left), whereas reasoning models typically include intermediate steps that reveal part of the thought process. (Note that many LLMs who have not been specifically developed for reasoning tasks can also provide intermediate reasoning steps in their answers.
大多数现代 LLM 具备基本推理能力,能回答诸如“如果一列火车以 60 英里/小时的速度行驶 3 小时,它走了多远?”之类的问题。因此,如今当我们提到推理模型时,通常指的是擅长更复杂推理任务的 LLM,例如解决谜题、谜语和数学证明。
Most modern LLMs are capable of basic reasoning and can answer questions like, "If a train is moving at 60 mph and travels for 3 hours, how far does it go?" So, today, when we refer to reasoning models, we typically mean LLMs that excel at more complex reasoning tasks, such as solving puzzles, riddles, and mathematical proofs.
此外,如今大多数被标榜为推理模型的大语言模型(LLM)都会在其响应中包含一个“思考”或“思维”过程。至于 LLM 是否真正“思考”以及如何“思考”,则是另一个讨论话题。
Additionally, most LLMs branded as reasoning models today include a "thought" or "thinking" process as part of their response. Whether and how an LLM actually "thinks" is a separate discussion.
推理模型中的中间步骤可以以两种方式出现。第一种,它们可能显式地包含在响应中,如上一图所示。第二种,一些推理 LLM(如 OpenAI 的 o1)会运行多次迭代,其中间步骤不向用户展示。
Intermediate steps in reasoning models can appear in two ways. First, they may be explicitly included in the response, as shown in the previous figure. Second, some reasoning LLMs, such as OpenAI's o1, run multiple iterations with intermediate steps that are not shown to the user.
[](https://substackcdn.com/image/fetch/$s_!DyRP!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F35712d0e-0f40-4855-8d81-4dcea94055ce_1538x810.png)
[](https://substackcdn.com/image/fetch/$s_!DyRP!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F35712d0e-0f40-4855-8d81-4dcea94055ce_1538x810.png)
_图 3:“推理”在两个不同层面使用:1)通过多个中间步骤处理输入并生成输出;2)在给用户的响应中提供某种推理。_
_Figure 3: "Reasoning" is used at two different levels: 1) processing the input and generating via multiple intermediate steps and 2) providing some sort of reasoning as part of the response to the user._
既然我们已经定义了推理模型,就可以进入更有趣的部分:如何构建和改进用于推理任务的 LLM。然而,在深入技术细节之前,重要的是考虑何时真正需要推理模型。
Now that we have defined reasoning models, we can move on to the more interesting part: how to build and improve LLMs for reasoning tasks. However, before diving into the technical details, it is important to consider when reasoning models are actually needed.
我们何时需要推理模型?推理模型旨在擅长复杂任务,如解决谜题、高级数学问题和具有挑战性的编码任务。然而,对于较简单的任务,如摘要、翻译或基于知识的问答,它们并非必需。事实上,对所有任务都使用推理模型可能效率低下且成本高昂。例如,推理模型通常使用成本更高、输出更冗长,有时还因“过度思考”而更容易出错。这里同样适用简单的规则:为任务使用合适的工具(或合适的 LLM 类型)。
When do we need a reasoning model? Reasoning models are designed to be good at complex tasks such as solving puzzles, advanced math problems, and challenging coding tasks. However, they are not necessary for simpler tasks like summarization, translation, or knowledge-based question answering. In fact, using reasoning models for everything can be inefficient and expensive. For instance, reasoning models are typically more expensive to use, more verbose, and sometimes more prone to errors due to "overthinking." Also here the simple rule applies: Use the right tool (or type of LLM) for the task.
下图总结了推理模型的主要优势和局限性。
The key strengths and limitations of reasoning models are summarized in the figure below.
[](https://substackcdn.com/image/fetch/$s_!lnf2!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F46dbe029-ab7d-4278-8dfe-7bc4af79a103_1352x524.png)
[](https://substackcdn.com/image/fetch/$s_!lnf2!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F46dbe029-ab7d-4278-8dfe-7bc4af79a103_1352x524.png)
_图 4:推理模型的主要优势与劣势。_
_Figure 4: The key strengths and weaknesses of reasoning models._
在下一节讨论构建和改进推理模型的四种主要方法之前,我想简要概述 DeepSeek R1 的流程,如 DeepSeek R1 技术报告所述。该报告既是一个有趣的案例研究,也是开发推理大语言模型的蓝图。
Before discussing four main approaches to building and improving reasoning models in the next section, I want to briefly outline the DeepSeek R1 pipeline, as described in the DeepSeek R1 technical report. This report serves as both an interesting case study and a blueprint for developing reasoning LLMs.
请注意,DeepSeek 并未发布单一的 R1 推理模型,而是引入了三个不同的变体:DeepSeek-R1-Zero、DeepSeek-R1 和 DeepSeek-R1-Distill。
Note that DeepSeek did not release a single R1 reasoning model but instead introduced three distinct variants: DeepSeek-R1-Zero, DeepSeek-R1, and DeepSeek-R1-Distill.
根据技术报告中的描述,我在下图中总结了这些模型的开发过程。
Based on the descriptions in the technical report, I have summarized the development process of these models in the diagram below.
[](https://substackcdn.com/image/fetch/$s_!z-dr!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb19df56-c5bf-4a0c-aafb-4629a39b13f5_1542x1166.png)
[](https://substackcdn.com/image/fetch/$s_!z-dr!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb19df56-c5bf-4a0c-aafb-4629a39b13f5_1542x1166.png)
图 5:DeepSeek R1 技术报告中讨论的 DeepSeek 三种不同推理模型的开发过程。
Figure 5: Development process of DeepSeek's three different reasoning models that are discussed in the DeepSeek R1 technical report.
接下来,我们简要回顾上图所示的过程。更多细节将在下一节中介绍,届时我们将讨论构建和改进推理模型的四种主要方法。
Next, let's briefly go over the process shown in the diagram above. More details will be covered in the next section, where we discuss the four main approaches to building and improving reasoning models.
(1) DeepSeek-R1-Zero:该模型基于 2024 年 12 月发布的 671B 预训练 DeepSeek-V3 基础模型。研究团队使用强化学习(RL)并采用两种奖励对其进行训练。这种方法被称为“冷启动”训练,因为它不包含监督微调(SFT)步骤,而该步骤通常是基于人类反馈的强化学习(RLHF)的一部分。
(1) DeepSeek-R1-Zero: This model is based on the 671B pre-trained DeepSeek-V3 base model released in December 2024. The research team trained it using reinforcement learning (RL) with two types of rewards. This approach is referred to as "cold start" training because it did not include a supervised fine-tuning (SFT) step, which is typically part of reinforcement learning with human feedback (RLHF).
(2) DeepSeek-R1:这是 DeepSeek 的旗舰推理模型,建立在 DeepSeek-R1-Zero 之上。团队通过额外的 SFT 阶段和进一步的 RL 训练对其进行了优化,改进了“冷启动”的 R1-Zero 模型。
(2) DeepSeek-R1: This is DeepSeek's flagship reasoning model, built upon DeepSeek-R1-Zero. The team further refined it with additional SFT stages and further RL training, improving upon the "cold-started" R1-Zero model.
(3) DeepSeek-R1-Distill*:利用前几步生成的 SFT 数据,DeepSeek 团队对 Qwen 和 Llama 模型进行了微调,以增强其推理能力。虽然这不是传统意义上的蒸馏,但该过程涉及在较大的 DeepSeek-R1 671B 模型的输出上训练较小的模型(Llama 8B 和 70B,以及 Qwen 1.5B–30B)。
(3) DeepSeek-R1-Distill*: Using the SFT data generated in the previous steps, the DeepSeek team fine-tuned Qwen and Llama models to enhance their reasoning abilities. While not distillation in the traditional sense, this process involved training smaller models (Llama 8B and 70B, and Qwen 1.5B–30B) on outputs from the larger DeepSeek-R1 671B model.
在本节中,我将概述当前用于增强大语言模型推理能力以及构建专用推理模型(如 DeepSeek-R1、OpenAI 的 o1 和 o3 等)的关键技术。
In this section, I will outline the key techniques currently used to enhance the reasoning capabilities of LLMs and to build specialized reasoning models such as DeepSeek-R1, OpenAI's o1 & o3, and others.
注意:o1 和 o3 的具体工作机制在 OpenAI 之外仍不为人知。然而,据传它们结合了推理和训练技术。
Note: The exact workings of o1 and o3 remain unknown outside of OpenAI. However, they are rumored to leverage a combination of both inference and training techniques.
提高 LLM 推理能力(或任何能力)的一种方法是推理时扩展。这个术语有多种含义,但在此上下文中,它指的是在推理期间增加计算资源以提高输出质量。
One way to improve an LLM's reasoning capabilities (or any capability in general) is inference-time scaling. This term can have multiple meanings, but in this context, it refers to increasing computational resources during inference to improve output quality.
一个粗略的类比是,人类在思考复杂问题时,如果给予更多时间,往往能产生更好的回答。类似地,我们可以应用一些技术来鼓励 LLM 在生成答案时“思考”更多。(尽管 LLM 是否真的“思考”是另一个讨论。)
A rough analogy is how humans tend to generate better responses when given more time to think through complex problems. Similarly, we can apply techniques that encourage the LLM to "think" more while generating an answer. (Although, whether LLMs actually "think" is a different discussion.)
推理时扩展的一种直接方法是巧妙的提示工程。一个经典的例子是_思维链(CoT)提示_,即在输入提示中包含“逐步思考”之类的短语。这鼓励模型生成中间推理步骤,而不是直接跳到最终答案,这通常(但并非总是)能在更复杂的问题上带来更准确的结果。(注意,对于简单的基于知识的问题,如“法国的首都是什么”,采用这种策略没有意义,这也是判断推理模型是否适用于给定输入查询的一个良好经验法则。)
One straightforward approach to inference-time scaling is clever prompt engineering. A classic example is _chain-of-thought (CoT) prompting_, where phrases like "think step by step" are included in the input prompt. This encourages the model to generate intermediate reasoning steps rather than jumping directly to the final answer, which can often (but not always) lead to more accurate results on more complex problems. (Note that it doesn't make sense to employ this strategy for simpler knowledge-based questions, like "What is the capital of France", which is again a good rule of thumb to find out whether a reasoning model makes sense on your given input query.)
[](https://substackcdn.com/image/fetch/$s_!VFAa!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F523eee5e-afb6-4019-a11b-e0a291d2c286_1600x419.png)
[](https://substackcdn.com/image/fetch/$s_!VFAa!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F523eee5e-afb6-4019-a11b-e0a291d2c286_1600x419.png)
_图 6:来自 2022 年论文《大型语言模型是零样本推理者》的经典 CoT 提示示例(https://arxiv.org/abs/2205.11916)。_
_Figure 6: An example of classic CoT prompting from the 2022 Large Language Models are Zero-Shot Reasoners paper (https://arxiv.org/abs/2205.11916)._
上述思维链方法可视为推理时扩展,因为它通过生成更多输出词元使得推理成本更高。
The aforementioned CoT approach can be seen as inference-time scaling because it makes inference more expensive through generating more output tokens.
推理时扩展的另一种方法是使用投票和搜索策略。一个简单的例子是多数投票,即让 LLM 生成多个答案,然后通过多数投票选择正确答案。类似地,我们可以使用束搜索和其他搜索算法来生成更好的响应。
Another approach to inference-time scaling is the use of voting and search strategies. One simple example is majority voting where we have the LLM generate multiple answers, and we select the correct answer by majority vote. Similarly, we can use beam search and other search algorithms to generate better responses.
我强烈推荐《Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters》这篇论文,我在之前的《2024 年值得关注的 AI 研究论文(第二部分)》文章(https://magazine.sebastianraschka.com/p/ai-research-papers-2024-part-2)中已经描述过,其中详细介绍了这些不同的策略。
I highly recommend the _Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters_ paper that I described in my previous Noteworthy AI Research Papers of 2024 (Part Two) article (https://magazine.sebastianraschka.com/p/ai-research-papers-2024-part-2) for more details on these different strategies.
[](https://substackcdn.com/image/fetch/$s_!YGJO!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5cb10e5a-738b-4c9e-ba65-5850d4793706_1600x919.png)
[](https://substackcdn.com/image/fetch/$s_!YGJO!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5cb10e5a-738b-4c9e-ba65-5850d4793706_1600x919.png)
_图 7:不同的基于搜索的方法依赖于基于过程奖励的模型来选择最佳答案。图注来自 LLM 测试时计算论文,https://arxiv.org/abs/2408.03314_
_Figure 7: Different search-based methods rely on a process-reward-based model to select the best answer. Annotated figure from the LLM Test-Time Compute paper, https://arxiv.org/abs/2408.03314_
DeepSeek R1 技术报告将常见的推理时扩展方法(如基于过程奖励模型和基于蒙特卡洛树搜索的方法)归类为“不成功的尝试”。这表明,除了 R1 模型自然倾向于生成更长响应(与 V3 基础模型相比,这构成了一种隐式的推理时扩展)之外,DeepSeek 并未显式使用这些技术。
The DeepSeek R1 technical report categorizes common inference-time scaling methods (such as Process Reward Model-based and Monte Carlo Tree Search-based approaches) under "unsuccessful attempts." This suggests that DeepSeek did not explicitly use these techniques beyond the R1 model's natural tendency to generate longer responses, which serves as an implicit form of inference-time scaling compared to the V3 base model.
然而,显式的推理时扩展通常在应用层实现,而非在 LLM 内部,因此 DeepSeek 仍可能在其应用中使用此类技术。
However, explicit inference-time scaling is often implemented at the application layer rather than within the LLM itself, so DeepSeek may still apply such techniques within their app.
我怀疑 OpenAI 的 o1 和 o3 模型使用了推理时扩展,这可以解释为什么它们相对于 GPT-4o 等模型更为昂贵。除了推理时扩展,o1 和 o3 很可能使用了与 DeepSeek R1 类似的强化学习流水线进行训练。关于强化学习的更多内容将在下面两节中介绍。
I suspect that OpenAI's o1 and o3 models use inference-time scaling, which would explain why they are relatively expensive compared to models like GPT-4o. In addition to inference-time scaling, o1 and o3 were likely trained using RL pipelines similar to those used for DeepSeek R1. More on reinforcement learning in the next two sections below.
DeepSeek R1 论文中我个人最欣赏的一点是,他们发现推理行为可以从纯强化学习(RL)中涌现。让我们更详细地探讨这意味着什么。
One of my personal highlights from the DeepSeek R1 paper is their discovery that reasoning emerges as a behavior from pure reinforcement learning (RL). Let's explore what this means in more detail.
如前所述,DeepSeek 开发了三种类型的 R1 模型。第一种是 DeepSeek-R1-Zero,它基于 DeepSeek-V3 基础模型构建,这是他们在 2024 年 12 月发布的标准预训练大语言模型。与典型的 RL 流程(在 RL 之前应用监督微调(SFT))不同,DeepSeek-R1-Zero 仅使用强化学习进行训练,没有初始的 SFT 阶段,如下图所示。
As outlined earlier, DeepSeek developed three types of R1 models. The first, DeepSeek-R1-Zero, was built on top of the DeepSeek-V3 base model, a standard pre-trained LLM they released in December 2024. Unlike typical RL pipelines, where supervised fine-tuning (SFT) is applied before RL, DeepSeek-R1-Zero was trained exclusively with reinforcement learning without an initial SFT stage as highlighted in the diagram below.
[](https://substackcdn.com/image/fetch/$s_!_9Z-!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa5bb6ecc-7e46-45fe-abff-1eb02e6b0e3a_1556x1162.png)
[](https://substackcdn.com/image/fetch/$s_!_9Z-!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa5bb6ecc-7e46-45fe-abff-1eb02e6b0e3a_1556x1162.png)
图 8:DeepSeek-R1-Zero 模型的开发过程。
Figure 8: The development process of DeepSeek-R1-Zero model.
不过,这种 RL 过程与常用的基于人类反馈的强化学习(RLHF)方法类似,后者通常用于对 LLM 进行偏好调整。(我在文章《LLM 训练:RLHF 及其替代方案》中更详细地介绍了 RLHF。)然而,如上所述,DeepSeek-R1-Zero 的关键区别在于,他们跳过了用于指令调整的监督微调(SFT)阶段。这就是他们称之为“纯”RL 的原因。(尽管在 LLM 语境下的 RL 与传统 RL 有很大不同,这将是另一个话题。)
Still, this RL process is similar to the commonly used RLHF approach, which is typically applied to preference-tune LLMs. (I covered RLHF in more detail in my article, _LLM Training: RLHF and Its Alternatives_.) However, as mentioned above, the key difference in _DeepSeek-R1-Zero_ is that they skipped the supervised fine-tuning (SFT) stage for instruction tuning. This is why they refer to it as "pure" RL. (Although, RL in the context of LLMs differs significantly from traditional RL, which is a topic for another time.)
对于奖励,他们没有使用基于人类偏好训练的奖励模型,而是采用了两种奖励:准确性奖励和格式奖励。
For rewards, instead of using a reward model trained on human preferences, they employed two types of rewards: an accuracy reward and a format reward.
* 准确性奖励使用 LeetCode 编译器验证编码答案,并使用确定性系统评估数学回答。
* The accuracy reward uses the LeetCode compiler to verify coding answers and a deterministic system to evaluate mathematical responses.
* 格式奖励依赖 LLM 评判器确保回答符合预期格式,例如将推理步骤放在 <think> 标签内。
* The format reward relies on an LLM judge to ensure responses follow the expected format, such as placing reasoning steps inside <think> tags.
令人惊讶的是,这种方法足以让 LLM 发展出基本的推理能力。研究人员观察到了一个“啊哈!”时刻,模型开始在其回答中生成推理轨迹,尽管并未被明确训练这样做,如下图所示。
Surprisingly, this approach was enough for the LLM to develop basic reasoning skills. The researchers observed an "Aha!" moment, where the model began generating reasoning traces as part of its responses despite not being explicitly trained to do so, as shown in the figure below.
[](https://substackcdn.com/image/fetch/$s_!Prn2!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30f8e37b-ba60-49d2-a95e-9c06b2033ee4_1600x1019.png)
[](https://substackcdn.com/image/fetch/$s_!Prn2!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30f8e37b-ba60-49d2-a95e-9c06b2033ee4_1600x1019.png)
_图 9:来自 DeepSeek R1 技术报告(https://arxiv.org/abs/2501.12948)的图,展示了“顿悟”时刻的出现。_
_Figure 9: A figure from the DeepSeek R1 technical report (https://arxiv.org/abs/2501.12948) showing the emergence of the "Aha" moment._
虽然 R1-Zero 并非性能顶尖的推理模型,但它通过生成中间的“思考”步骤展示了推理能力,如上图所示。这证实了使用纯强化学习开发推理模型是可行的,而 DeepSeek 团队是首个展示(或至少发表)这一方法的团队。
While R1-Zero is not a top-performing reasoning model, it does demonstrate reasoning capabilities by generating intermediate "thinking" steps, as shown in the figure above. This confirms that it is possible to develop a reasoning model using pure RL, and the DeepSeek team was the first to demonstrate (or at least publish) this approach.
《Ahead of AI》是一份由读者支持的出版物。要接收新文章并支持我的工作,请考虑成为免费或付费订阅者。
Ahead of AI is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.
接下来,我们来看 DeepSeek 旗舰推理模型 DeepSeek-R1 的开发过程,它可作为构建推理模型的蓝图。该模型在 DeepSeek-R1-Zero 的基础上,通过引入额外的监督微调(SFT)和强化学习(RL)来提升其推理性能。
Next, let's look at the development of DeepSeek-R1, DeepSeek's flagship reasoning model, which serves as a blueprint for building reasoning models. This model improves upon DeepSeek-R1-Zero by incorporating additional supervised fine-tuning (SFT) and reinforcement learning (RL) to improve its reasoning performance.
注意,在 RL 之前加入 SFT 阶段实际上是常见的做法,如标准 RLHF 流程所示。OpenAI 的 o1 很可能采用了类似的方法。
Note that it is actually common to include an SFT stage before RL, as seen in the standard RLHF pipeline. OpenAI's o1 was likely developed using a similar approach.
[](https://substackcdn.com/image/fetch/$s_!19pK!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdf7f99f0-d154-49e5-b60a-4d148e0a61be_1548x1154.png)
[](https://substackcdn.com/image/fetch/$s_!19pK!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdf7f99f0-d154-49e5-b60a-4d148e0a61be_1548x1154.png)
图 10:DeepSeek-R1 模型的开发流程。
Figure 10: The development process of DeepSeek-R1 model.
如上图所示,DeepSeek 团队使用 DeepSeek-R1-Zero 生成了他们所谓的“冷启动”SFT 数据。“冷启动”一词指的是这些数据由 DeepSeek-R1-Zero 生成,而该模型本身并未经过任何监督微调(SFT)数据的训练。
As shown in the diagram above, the DeepSeek team used DeepSeek-R1-Zero to generate what they call "cold-start" SFT data. The term "cold start" refers to the fact that this data was produced by DeepSeek-R1-Zero, which itself had not been trained on any supervised fine-tuning (SFT) data.
利用这些冷启动 SFT 数据,DeepSeek 随后通过指令微调训练模型,并紧接着进行了另一阶段的强化学习(RL)。该 RL 阶段保留了 DeepSeek-R1-Zero 的 RL 过程中使用的相同准确率和格式奖励。然而,他们增加了一致性奖励以防止语言混合,即模型在回答中切换多种语言的现象。
Using this cold-start SFT data, DeepSeek then trained the model via instruction fine-tuning, followed by another reinforcement learning (RL) stage. This RL stage retained the same accuracy and format rewards used in DeepSeek-R1-Zero’s RL process. However, they added a consistency reward to prevent language mixing, which occurs when the model switches between multiple languages within a response.
RL 阶段之后又进行了一轮 SFT 数据收集。在此阶段,使用最新的模型检查点生成了 600K 个思维链(CoT)SFT 示例,同时使用 DeepSeek-V3 基础模型创建了额外的 200K 个基于知识的 SFT 示例。
The RL stage was followed by another round of SFT data collection. In this phase, the most recent model checkpoint was used to generate 600K Chain-of-Thought (CoT) SFT examples, while an additional 200K knowledge-based SFT examples were created using the DeepSeek-V3 base model.
这些 600K + 200K 的 SFT 样本随后用于对 DeepSeek-V3 基础模型进行指令微调,之后再进行最后一轮 RL。在此阶段,他们再次使用基于规则的方法对数学和编码问题提供准确率奖励,而其他类型的问题则使用人类偏好标签。总而言之,这与常规的基于人类反馈的强化学习(RLHF)非常相似,只是 SFT 数据包含(更多)CoT 示例,并且 RL 除了基于人类偏好的奖励外,还具有可验证的奖励。
These 600K + 200K SFT samples were then used for instruction-finetuning DeepSeek-V3 base before following up with a final round of RL. In this stage, they again used rule-based methods for accuracy rewards for math and coding questions, while human preference labels used for other question types. All in all, this is very similar to regular RLHF except that the SFT data contains (more) CoT examples. And the RL has verifiable rewards in addition to human preference-based rewards.
最终模型 DeepSeek-R1 相比 DeepSeek-R1-Zero 有显著的性能提升,这得益于额外的 SFT 和 RL 阶段,如下表所示。
The final model, DeepSeek-R1 has a noticeable performance boost over DeepSeek-R1-Zero thanks to the additional SFT and RL stages, as shown in the table below.
[](https://substackcdn.com/image/fetch/$s_!22Cm!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff7f73f16-db4e-4047-89b0-823f16cefb33_1556x490.png)
[](https://substackcdn.com/image/fetch/$s_!22Cm!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff7f73f16-db4e-4047-89b0-823f16cefb33_1556x490.png)
图 11:OpenAI O1 与 DeepSeek R1 模型的基准对比。该图来自 DeepSeek-R1 技术报告(https://arxiv.org/abs/2501.12948),并附有注释。
Figure 11: Benchmark comparison of OpenAI O1 and DeepSeek R1 models. Annotated figure from the DeepSeek-R1 technical report (https://arxiv.org/abs/2501.12948).
到目前为止,我们已经介绍了构建和改进推理模型的三种关键方法:
So far, we have covered three key approaches to building and improving reasoning models:
1. 推理时扩展(inference-time scaling),一种无需训练或修改底层模型即可提升推理能力的技术。
1. Inference-time scaling, a technique that improves reasoning capabilities without training or otherwise modifying the underlying model.
2. 纯强化学习(RL),如 DeepSeek-R1-Zero 所示,它表明推理能力可以在没有监督微调的情况下作为学习行为涌现。
2. Pure reinforcement learning (RL) as in DeepSeek-R1-Zero, which showed that reasoning can emerge as a learned behavior without supervised fine-tuning.
3. 监督微调(SFT)加强化学习,这催生了 DeepSeek 的旗舰推理模型 DeepSeek-R1。
3. Supervised fine-tuning (SFT) plus RL, which led to DeepSeek-R1, DeepSeek’s flagship reasoning model.
令人惊讶的是,DeepSeek 还发布了通过他们称为“蒸馏”的过程训练的较小模型。然而,在 LLM 的背景下,蒸馏并不一定遵循深度学习中经典的知识蒸馏方法。传统上,在知识蒸馏中(如我的《机器学习问答与 AI》一书第 6 章简要描述的),较小的学生模型在较大教师模型的 logits 和目标数据集上进行训练。
Surprisingly, DeepSeek also released smaller models trained via a process they call _distillation_. However, in the context of LLMs, distillation does not necessarily follow the classical knowledge distillation approach used in deep learning. Traditionally, in knowledge distillation (as briefly described in Chapter 6 of my Machine Learning Q and AI book), a smaller student model is trained on both the logits of a larger teacher model and a target dataset.
相反,这里的蒸馏指的是在由较大 LLM 生成的 SFT 数据集上,对较小的 LLM(如 Llama 8B 和 70B,以及 Qwen 2.5 系列模型(0.5B 至 32B))进行指令微调。具体来说,这些较大的 LLM 是 DeepSeek-V3 和 DeepSeek-R1 的一个中间检查点。事实上,用于此蒸馏过程的 SFT 数据与上一节所述的用于训练 DeepSeek-R1 的数据集相同。
Instead, here distillation refers to instruction fine-tuning smaller LLMs, such as Llama 8B and 70B and Qwen 2.5 models (0.5B to 32B), on an SFT dataset generated by larger LLMs. Specifically, these larger LLMs are DeepSeek-V3 and an intermediate checkpoint of DeepSeek-R1. In fact, the SFT data used for this distillation process is the same dataset that was used to train DeepSeek-R1, as described in the previous section.
为了澄清这一过程,我在下图中突出显示了蒸馏部分。
To clarify this process, I have highlighted the distillation portion in the diagram below.
[](https://substackcdn.com/image/fetch/$s_!xUjE!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7db7c46b-fe67-49f4-9f65-b0e7b7e5ac08_1444x1174.png)
[](https://substackcdn.com/image/fetch/$s_!xUjE!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7db7c46b-fe67-49f4-9f65-b0e7b7e5ac08_1444x1174.png)
图 12:DeepSeek-R1-Distill 模型的开发过程。
Figure 12: The development process of DeepSeek-R1-Distill models.
他们为什么开发这些蒸馏模型?在我看来,有两个关键原因:
Why did they develop these distilled models? In my opinion, there are two key reasons:
1. 较小的模型更高效。这意味着它们运行成本更低,而且可以在较低端的硬件上运行,这使得它们对许多研究人员和像我这样的爱好者特别有吸引力。
1. Smaller models are more efficient. This means they are cheaper to run, but they also can run on lower-end hardware, which makes these especially interesting for many researchers and tinkerers like me.
2. 纯 SFT 的案例研究。这些蒸馏模型作为一个有趣的基准,展示了在没有强化学习的情况下,纯监督微调(SFT)能将模型提升到何种程度。
2. A case study in pure SFT. These distilled models serve as an interesting benchmark, showing how far pure supervised fine-tuning (SFT) can take a model without reinforcement learning.
下表将这些蒸馏模型与其他流行模型以及 DeepSeek-R1-Zero 和 DeepSeek-R1 的性能进行了比较。
The table below compares the performance of these distilled models against other popular models, as well as DeepSeek-R1-Zero and DeepSeek-R1.
[](https://substackcdn.com/image/fetch/$s_!XwZe!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Febc749fb-6a79-483f-bcda-b219f284bc09_1168x604.png)
[](https://substackcdn.com/image/fetch/$s_!XwZe!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Febc749fb-6a79-483f-bcda-b219f284bc09_1168x604.png)
图 13:蒸馏模型与非蒸馏模型的基准比较。图来自 DeepSeek-R1 技术报告(https://arxiv.org/abs/2501.12948),已添加注释。
Figure 13: Benchmark comparison of distilled versus non-distilled models. Annotated figure from the DeepSeek-R1 technical report (https://arxiv.org/abs/2501.12948).
正如我们所见,蒸馏模型明显弱于 DeepSeek-R1,但相对于 DeepSeek-R1-Zero,它们却出奇地强大,尽管规模小了若干个数量级。同样值得注意的是,这些模型与 o1-mini 相比表现如何(我怀疑 o1-mini 本身可能就是 o1 的类似蒸馏版本)。
As we can see, the distilled models are noticeably weaker than DeepSeek-R1, but they are surprisingly strong relative to DeepSeek-R1-Zero, despite being orders of magnitude smaller. It's also interesting to note how well these models perform compared to o1-mini (I suspect o1-mini itself might be a similarly distilled version of o1).
在结束本节之前,还有一个有趣的比较值得一提。DeepSeek 团队测试了在 DeepSeek-R1-Zero 中出现的涌现推理行为是否也能在较小的模型中出现。为此,他们直接将 DeepSeek-R1-Zero 的纯强化学习方法应用于 Qwen-32B。
Before wrapping up this section with a conclusion, there’s one more interesting comparison worth mentioning. The DeepSeek team tested whether the emergent reasoning behavior seen in DeepSeek-R1-Zero could also appear in smaller models. To investigate this, they applied the same pure RL approach from DeepSeek-R1-Zero directly to Qwen-32B.
该实验的结果总结在下表中,其中 QwQ-32B-Preview 作为基于 Qwen 2.5 32B 的参考推理模型,由 Qwen 团队开发(我认为训练细节从未公开)。这一比较为纯强化学习是否能在远小于 DeepSeek-R1-Zero 的模型中诱导推理能力提供了额外的见解。
The results of this experiment are summarized in the table below, where QwQ-32B-Preview serves as a reference reasoning model based on Qwen 2.5 32B developed by the Qwen team (I think the training details were never disclosed). This comparison provides some additional insights into whether pure RL alone can induce reasoning capabilities in models much smaller than DeepSeek-R1-Zero.
[](https://substackcdn.com/image/fetch/$s_!5_5L!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F05514c9f-eb04-496b-bd98-bb4710c65b14_1448x408.png)
[](https://substackcdn.com/image/fetch/$s_!5_5L!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F05514c9f-eb04-496b-bd98-bb4710c65b14_1448x408.png)
图 14:在较小的 32B 模型上进行蒸馏和强化学习的基准比较。来自 DeepSeek-R1 技术报告的注释图(https://arxiv.org/abs/2501.12948)。
Figure 14: Benchmark comparison distillation and RL on a smaller 32B model. Annotated figure from the DeepSeek-R1 technical report (https://arxiv.org/abs/2501.12948).
有趣的是,结果表明对于较小的模型,蒸馏远比纯强化学习有效。这与以下观点一致:仅靠强化学习可能不足以在如此规模的模型中引发强大的推理能力,而在高质量推理数据上进行监督微调(SFT)对于小模型而言可能是更有效的策略。
Interestingly, the results suggest that distillation is far more effective than pure RL for smaller models. This aligns with the idea that RL alone may not be sufficient to induce strong reasoning abilities in models of this scale, whereas SFT on high-quality reasoning data can be a more effective strategy when working with small models.
为了完整性,在表格中看到额外的对比将会很有用:
For completeness, it would have been useful to see additional comparisons in the table:
1. 使用 SFT + RL 训练的 Qwen-32B,类似于 DeepSeek-R1 的开发方式。这将有助于确定与纯 RL 和纯 SFT 相比,当 RL 与 SFT 结合时能带来多大的提升。
1. Qwen-32B trained with SFT + RL, similar to how DeepSeek-R1 was developed. This would help determine how much improvement can be made, compared to pure RL and pure SFT, when RL is combined with SFT.
2. 使用纯 SFT 训练的 DeepSeek-V3,类似于蒸馏模型的创建方式。这将允许直接比较,以了解 RL + SFT 相对于纯 SFT 的有效性。
2. DeepSeek-V3 trained with pure SFT, similar to how the distilled models were created. This would allow for a direct comparison to see how effective RL + SFT is over pure SFT.
《AI 前沿》是一份由读者支持的出版物。要接收新文章并支持我的工作,请考虑成为免费或付费订阅者。
Ahead of AI is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.
在本节中,我们探讨了构建和改进推理模型的四种不同策略:
In this section, we explored four different strategies for building and improving reasoning models:
1. 推理时扩展(inference-time scaling)不需要额外训练,但会增加推理成本,随着用户数量或查询量的增长,大规模部署会变得更加昂贵。尽管如此,对于提升已有强模型的性能,它仍然是一个无需多虑的选择。我强烈怀疑 o1 利用了推理时扩展,这有助于解释为什么它在每 token 成本上比 DeepSeek-R1 更贵。
1. Inference-time scaling requires no additional training but increases inference costs, making large-scale deployment more expensive as the number of users or query volume grows. Still, it remains a no-brainer for improving the performance of already strong models. I strongly suspect that o1 leverages inference-time scaling, which helps explain why it is more expensive on a per-token basis compared to DeepSeek-R1.
2. 纯强化学习(RL)在研究方面很有趣,因为它提供了对推理作为涌现行为的洞察。然而,在实际的模型开发中,RL + SFT 是首选方法,因为它能带来更强的推理模型。我也强烈怀疑 o1 使用了 RL + SFT 进行训练。更准确地说,我认为 o1 从一个比 DeepSeek-R1 更弱、更小的基础模型开始,但通过 RL + SFT 和推理时扩展进行了补偿。
2. Pure RL is interesting for research purposes because it provides insights into reasoning as an emergent behavior. However, in practical model development, RL + SFT is the preferred approach as it leads to stronger reasoning models. I strongly suspect that o1 was trained using RL + SFT as well. More precisely, I believe o1 starts from a weaker, smaller base model than DeepSeek-R1 but compensates with RL + SFT and inference-time scaling.
3. 如上所述,RL + SFT 是构建高性能推理模型的关键方法。DeepSeek-R1 是一个很好的蓝图,展示了如何做到这一点。
3. As mentioned above, RL + SFT is the key approach for building high-performance reasoning models. DeepSeek-R1 is a nice blueprint showing how this can be done.
4. 蒸馏(distillation)是一种有吸引力的方法,特别是用于创建更小、更高效的模型。然而,其局限性在于蒸馏不会推动创新,也不会产生下一代推理模型。例如,蒸馏总是依赖于一个现有的、更强的模型来生成监督微调(SFT)数据。
4. Distillation is an attractive approach, especially for creating smaller, more efficient models. However, the limitation is that distillation does not drive innovation or produce the next generation of reasoning models. For instance, distillation always depends on an existing, stronger model to generate the supervised fine-tuning (SFT) data.
我期待的下一个有趣方向是将强化学习与监督微调(方法 3)与推理时扩展(方法 1)相结合。这很可能就是 OpenAI o1 正在做的事情,只不过它可能基于比 DeepSeek-R1 更弱的基座模型,这解释了为什么 DeepSeek-R1 在推理时保持相对廉价的同时表现如此出色。
One interesting aspect I expect to see next is to combine RL + SFT (approach 3) with inference-time scaling (approach 1). This is likely what OpenAI o1 is doing, except it's probably based on a weaker base model than DeepSeek-R1, which explains why DeepSeek-R1 performs so well while remaining relatively cheap at inference time.
最近几周,许多人询问我对 DeepSeek-R1 模型的看法。简而言之,我认为这是一项了不起的成就。作为一名研究工程师,我尤其欣赏其详细的技术报告,其中提供了我能够学习的方法论见解。
In recent weeks, many people have asked for my thoughts on the DeepSeek-R1 models. In short, I think they are an awesome achievement. As a research engineer, I particularly appreciate the detailed technical report, which provides insights into their methodology that I can learn from.
最引人入胜的收获之一是,推理行为如何从纯粹的强化学习中涌现。而且,DeepSeek 以宽松的开源 MIT 许可证开源其模型,这比 Meta 的 Llama 模型的限制更少,令人印象深刻。
One of the most fascinating takeaways is how reasoning emerged as a behavior from pure RL. And it's impressive that DeepSeek has open-sourced their models under a permissive open-source MIT license, which has even fewer restrictions than Meta's Llama models.
DeepSeek-R1 比 o1 更好吗?我认为它们大致处于同一水平。然而,突出的一点是 DeepSeek-R1 在推理时更高效。这表明 DeepSeek 可能在训练过程中投入更多,而 OpenAI 可能更多依赖 o1 的推理时扩展。
Is DeepSeek-R1 better than o1? I’d say it’s roughly in the same ballpark. However, what stands out is that DeepSeek-R1 is more efficient at inference time. This suggests that DeepSeek likely invested more heavily in the training process, while OpenAI may have relied more on inference-time scaling for o1.
尽管如此,直接比较 o1 和 DeepSeek-R1 是困难的,因为 OpenAI 没有透露太多关于 o1 的信息。例如,我们不知道:
That said, it's difficult to compare o1 and DeepSeek-R1 directly because OpenAI has not disclosed much about o1. For instance, we don’t know:
* o1 是否也是混合专家(MoE)模型?
* Is o1 also a Mixture of Experts (MoE)?
o1 会不会只是 GPT-4o 的一个略微改进版本,仅需最少的强化学习(RL)和微调(SFT),并大量依赖推理时的扩展?
Could o1 just be a slightly refined version of GPT-4o with minimal RL + SFT and only extensive inference-time scaling?
在不了解这些细节的情况下,直接比较仍然是苹果与橘子的比较。
Without knowing these details, a direct comparison remains an apples-to-oranges comparison.
另一个讨论点是开发 DeepSeek-R1 的成本。有人提到约 600 万美元的训练成本,但他们很可能混淆了 DeepSeek-V3(去年 12 月发布的基础模型)和 DeepSeek-R1。
Another point of discussion has been the cost of developing DeepSeek-R1. Some have mentioned a ~$6 million training cost, but they likely conflated DeepSeek-V3 (the base model released in December last year) and DeepSeek-R1.
600 万美元的估算基于假设每 GPU 小时 2 美元,以及 DeepSeek-V3 最终训练运行所需的 GPU 小时数,这一数字最初是在 2024 年 12 月讨论的。
The $6 million estimate is based on an assumed $2 per GPU hour and the number of GPU hours required for the final training run of DeepSeek-V3, which was originally discussed back in December 2024.
然而,DeepSeek 团队从未披露 R1 的确切 GPU 小时数或开发成本,因此任何成本估算都纯属猜测。
However, the DeepSeek team has never disclosed the exact GPU hours or development cost for R1, so any cost estimates remain pure speculation.
无论如何,最终,DeepSeek-R1 是开放权重推理模型的一个重要里程碑,其在推理时的效率使其成为 OpenAI o1 的一个有趣替代方案。
Either way, ultimately, DeepSeek-R1 is a major milestone in open-weight reasoning models, and its efficiency at inference time makes it an interesting alternative to OpenAI’s o1.
开发一个达到 DeepSeek-R1 水平的推理模型,即使从像 DeepSeek-V3 这样的开放权重基础模型开始,也可能需要数十万到数百万美元。这对于预算有限的研究人员或工程师来说,可能会感到沮丧。
Developing a DeepSeek-R1-level reasoning model likely requires hundreds of thousands to millions of dollars, even when starting with an open-weight base model like DeepSeek-V3. This can feel discouraging for researchers or engineers working with limited budgets.
好消息是:蒸馏(Distillation)可以大有作为
The good news: Distillation can go a long way
幸运的是,模型蒸馏提供了一种更具成本效益的替代方案。DeepSeek 团队通过他们的 R1 蒸馏模型证明了这一点,这些模型尽管比 DeepSeek-R1 小得多,却实现了令人惊讶的强大推理性能。然而,即使这种方法也并非完全便宜。他们的蒸馏过程使用了 80 万条 SFT 样本,这需要大量的算力。
Fortunately, model distillation offers a more cost-effective alternative. The DeepSeek team demonstrated this with their R1-distilled models, which achieve surprisingly strong reasoning performance despite being significantly smaller than DeepSeek-R1. However, even this approach isn’t entirely cheap. Their distillation process used 800K SFT samples, which requires substantial compute.
有趣的是,就在 DeepSeek-R1 发布前几天,我偶然看到一篇关于 Sky-T1 的文章,这是一个引人入胜的项目,一个小团队仅用 1.7 万条 SFT 样本就训练了一个开放权重的 32B 模型。总成本?仅 450 美元,这比大多数 AI 会议的注册费还要少。
Interestingly, just a few days before DeepSeek-R1 was released, I came across an article about Sky-T1, a fascinating project where a small team trained an open-weight 32B model using only 17K SFT samples. The total cost? Just $450, which is less than the registration fee for most AI conferences.
这个例子表明,虽然大规模训练仍然昂贵,但小规模、有针对性的微调工作仍然可以以极低的成本取得令人印象深刻的结果。
This example highlights that while large-scale training remains expensive, smaller, targeted fine-tuning efforts can still yield impressive results at a fraction of the cost.
[](https://substackcdn.com/image/fetch/$s_!Y8HI!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8865a313-2326-4f07-a6dc-72cc94cb2ebe_1364x570.png)
[](https://substackcdn.com/image/fetch/$s_!Y8HI!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8865a313-2326-4f07-a6dc-72cc94cb2ebe_1364x570.png)
图 15:来自文章《Sky-T1:在 450 美元内训练你自己的 O1 预览模型》的图,https://novasky-ai.github.io/posts/sky-t1/
Figure 15: Figure from the "Sky-T1: Train your own O1 preview model within $450" article, https://novasky-ai.github.io/posts/sky-t1/
根据他们的基准测试,Sky-T1 的表现大致与 o1 相当,考虑到其低廉的训练成本,这令人印象深刻。
According to their benchmarks, Sky-T1 performs roughly on par with o1, which is impressive given its low training cost.
虽然 Sky-T1 专注于模型蒸馏,但我也在“纯强化学习”领域发现了一些有趣的工作。一个显著的例子是 TinyZero,一个 3B 参数的模型,它复现了 DeepSeek-R1-Zero 的方法(附带说明:训练成本不到 30 美元)。
While Sky-T1 focused on model distillation, I also came across some interesting work in the "pure RL" space. One notable example is TinyZero, a 3B parameter model that replicates the DeepSeek-R1-Zero approach (side note: it costs less than $30 to train).
令人惊讶的是,即使只有 3B 参数,TinyZero 也展现出一些涌现的自我验证能力,这支持了推理可以通过纯强化学习涌现的观点,即使在小型模型中也是如此。
Surprisingly, even at just 3B parameters, TinyZero exhibits some emergent self-verification abilities, which supports the idea that reasoning can emerge through pure RL, even in small models.
TinyZero 仓库提到一份研究报告仍在撰写中,我一定会密切关注后续细节。
The TinyZero repository mentions that a research report is still work in progress, and I’ll definitely be keeping an eye out for further details.
图 16:来自 TinyZero 仓库(https://github.com/Jiayi-Pan/TinyZero)的图,显示模型能够进行自我验证。(如果能同时看到基础模型的响应作为对比,将会很有趣。)
Figure 16: A figure from the TinyZero repository (https://github.com/Jiayi-Pan/TinyZero) showing that the model is capable of self-verification. (It would have been interesting to see the response of the base model in comparison.)
上述两个项目表明,即使在有限的预算下,也能开展有趣的推理模型研究。虽然这两种方法都复现了 DeepSeek-R1 的方法,一个专注于纯强化学习(TinyZero),另一个专注于纯监督微调(Sky-T1),但探索这些想法如何进一步扩展将是引人入胜的。
The two projects mentioned above demonstrate that interesting work on reasoning models is possible even with limited budgets. While both approaches replicate methods from DeepSeek-R1, one focusing on pure RL (TinyZero) and the other on pure SFT (Sky-T1), it would be fascinating to explore how these ideas can be extended further.
去年我遇到的一个特别有趣的方法描述在论文《O1 复现之旅:战略进展报告——第一部分》(https://arxiv.org/abs/2410.18982)中。尽管标题如此,该论文实际上并未复现 o1。相反,它引入了一种不同的方式来改进蒸馏(纯监督微调)过程。
One particularly interesting approach I came across last year is described in the paper _O1 Replication Journey: A Strategic Progress Report – Part 1_ (https://arxiv.org/abs/2410.18982). Despite its title, the paper does not actually replicate o1. Instead, it introduces a different way to improve the distillation (pure SFT) process.
本文的核心思想是“旅程学习”(journey learning),作为“捷径学习”(shortcut learning)的替代方案。
The key idea in the paper is "journey learning" as an alternative to "shortcut learning."
* 捷径学习指的是指令微调中的传统方法,即仅使用正确的解题路径来训练模型。
* Shortcut learning refers to the traditional approach in instruction fine-tuning, where models are trained using only correct solution paths.
* 相比之下,旅程学习还包括错误的解题路径,使模型能够从错误中学习。
* Journey learning, on the other hand, also includes incorrect solution paths, allowing the model to learn from mistakes.
这种方法与 TinyZero 纯强化学习训练中观察到的自我验证能力有一定关联,但它侧重于完全通过 SFT 来改进模型。通过让模型接触错误的推理路径及其修正,旅程学习可能还会增强自我纠正能力,从而可能使推理模型更加可靠。
This approach is kind of related to the self-verification abilities observed in TinyZero’s pure RL training, but it focuses on improving the model entirely through SFT. By exposing the model to incorrect reasoning paths and their corrections, journey learning may also reinforce self-correction abilities, potentially making reasoning models more reliable this way.
[](https://substackcdn.com/image/fetch/$s_!TxCO!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a0bfcd0-6d93-4c91-a0d6-28178839b7cf_1492x724.png)
[](https://substackcdn.com/image/fetch/$s_!TxCO!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a0bfcd0-6d93-4c91-a0d6-28178839b7cf_1492x724.png)
图 17:旅程学习(Journey learning)与传统捷径学习(shortcut learning)不同,它在 SFT 数据中包含了错误的解题路径。该图为注释图,来自《O1 复制之旅:战略进展报告——第一部分》(https://arxiv.org/abs/2410.18982)。
Figure 17: Journey learning, as opposed to traditional shortcut learning, includes wrong solution paths in the SFT data. Annotated figure from the O1 Replication Journey: A Strategic Progress Report – Part 1 (https://arxiv.org/abs/2410.18982)
这可能是未来工作中一个令人兴奋的方向,尤其是在低预算推理模型开发中,因为基于强化学习的方法在计算上可能不切实际。
This could be an exciting direction for future work, particularly for low-budget reasoning model development, where RL-based approaches may be computationally impractical.
总之,目前推理模型领域涌现了大量有趣的工作,我相信在接下来的几个月里我们还会看到更多令人兴奋的进展!
Anyways, a lot of interesting work is currently happening on the reasoning model front, and I'm sure we will see a lot more exciting work in the upcoming months!
_本杂志是个人热情项目。若您愿意支持我,请考虑购买我的《构建大语言模型(从零开始)》一书。(我确信您会从这本书中获益良多,因为它以其他地方难以企及的详细程度解释了 LLM 的工作原理。)_
_This magazine is a personal passion project. For those who wish to support me, please consider purchasing a copy of my Build a Large Language Model (From Scratch) book. (I am confident that you'll get lots out of this book as it explains how LLMs work in a level of detail that is not found anywhere else.)_
[](https://substackcdn.com/image/fetch/$s_!woQp!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea1152a0-18d9-4a8a-9398-c6b1ca67726a_1600x900.png)
[](https://substackcdn.com/image/fetch/$s_!woQp!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea1152a0-18d9-4a8a-9398-c6b1ca67726a_1600x900.png)
《从头构建大语言模型》现已在亚马逊上架
Build a Large Language Model (From Scratch) now available on Amazon
_如果您读过这本书并且有几分钟空闲,我将非常感激您能写一篇简短的书评。这对我们作者帮助很大!_
_If you read the book and have a few minutes to spare, I'd really appreciate a brief review. It helps us authors a lot!_
您的支持意义重大!谢谢!
Your support means a great deal! Thank you!