Controlling Reasoning Effort in LLMs
打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→自 OpenAI 发布 o1 以来,已经过去了近两年,o1 是一个普及了基于 LLM 的推理模型理念的模型。大约四个月后,DeepSeek-R1 紧随其后,并提供了使用可验证奖励的强化学习(RLVR)配方来训练此类推理模型的细节。上周,OpenAI 发布了 GPT-5.6 模型系列。该系列包含三种尺寸,每种尺寸大约有五六种推理努力设置。图 1:GPT 5.6 Sol 模型在不同推理努力设置下的表现。(Ultra 的基准数字目前尚不可用,但应与 Max 相对相似,因为它使用类似的努力水平,但通过四个子代理加速工作。)
It has been almost two years since OpenAI released o1, a model that popularized the idea of LLM-based reasoning models. DeepSeek-R1 followed about four months later, together with details of a reinforcement learning with verifiable rewards (RLVR) recipe to train such reasoning models. Last week, OpenAI released the GPT-5.6 model family. It comes in three sizes, each with roughly five or six reasoning-effort settings. Figure 1: The GPT 5.6 Sol model with different reasoning effort settings. (Benchmark numbers for Ultra are currently not available but should be relatively similar to Max, since it uses a similar effort level but accelerates the work with four subagents.)
距离 OpenAI 发布 o1 已经将近两年,该模型普及了基于 LLM 的推理模型的概念。大约四个月后,DeepSeek-R1 紧随其后,并提供了使用可验证奖励的强化学习(RLVR)方法来训练此类推理模型的细节。
It has been almost two years since OpenAI released o1, a model that popularized the idea of LLM-based reasoning models. DeepSeek-R1 followed about four months later, together with details of a reinforcement learning with verifiable rewards (RLVR) recipe to train such reasoning models.
上周,OpenAI 发布了 GPT-5.6 模型系列。该系列包含三种尺寸,每种尺寸大约有五到六种推理努力设置。
Last week, OpenAI released the GPT-5.6 model family. It comes in three sizes, each with roughly five or six reasoning-effort settings.
图 1:GPT 5.6 Sol 模型在不同推理努力设置下的表现。(Ultra 的基准测试数据目前尚未公布,但应与 Max 相对接近,因为它使用类似的努力水平,但通过四个子智能体加速工作。)
Figure 1: The GPT 5.6 Sol model with different reasoning effort settings. (Benchmark numbers for Ultra are currently not available but should be relatively similar to Max, since it uses a similar effort level but accelerates the work with four subagents.)
所以,是的,推理模型已经站稳脚跟。它们已成为现代模型发布的标准组成部分。
So yes, reasoning models are here to stay. They have become a standard part of modern model releases.
过去,我介绍了推理模型的方法论(理解推理 LLM)以及相关研究论文(LLM 推理的强化学习现状和 LLM 推理模型推理现状)。我甚至写了一本全新的 440 页关于如何开发推理模型的书《构建推理模型(从零开始)》。
In the past, I covered the methodology of reasoning models (Understanding Reasoning LLMs) as well as relevant research papers (The State of Reinforcement Learning for LLM Reasoning and The State of LLM Reasoning Model Inference). And I even wrote a whole new 440-page book on how to develop reasoning models, Build A Reasoning Model (From Scratch).
图 2:我的新书《构建推理模型(从零开始)》,全彩!
Figure 2: My new Build A Reasoning Model (From Scratch) book. In color!
这些资源专注于将传统 LLM 转变为推理模型。现在,在这篇文章中,我想重点解释如何开发一个具有多种努力模式的推理模型,类似于本文开头图中所示。
These resources have focused on turning a conventional LLM into a reasoning model. Now, in this article, I want to focus on and explain how to develop a reasoning model that has multiple effort modes, similar to what’s shown in the figure at the beginning of this article.
别担心,这篇文章可以独立阅读。不过,上述资源可能也很有趣且有用。
No worries, this article can be read as a standalone article. However, the aforementioned resources may be interesting and useful.
在讨论几乎任何机器学习或 AI 技术或子领域时,一个重要的教训是:我们通常不应从字面上理解技术术语。例如,机器学习与 AI 中的(人工)神经网络并非真的像人脑这样的生物神经网络那样运作。
When talking about pretty much any machine learning or AI technique or subfield, the one lesson is that we usually shouldn’t take technical terms “literally”. For example, an (artificial) neural network in machine learning and AI doesn’t literally work like a biological neural network like the human brain.
同样,当我们谈论“推理模型”时,不应期望这些模型真的像人类一样推理。在 AI 和 LLM 研究的语境中,“推理模型”指的是输出中间推理轨迹的模型,这种轨迹类似于逐步解答问题或任务的中间响应。
Similarly, when talking about “reasoning models”, we shouldn’t expect that these models literally reason like us humans. In the context of AI and LLM research, “reasoning model” means a model that outputs an intermediate reasoning trace, which is like an intermediate response that works through a question or task step by step.
通过一个例子来解释可能最容易理解。
It’s probably easiest to explain this by showing an example.
图 3:传统 LLM 回答(左)与推理模型回答(右)的示意图。
Figure 3: Illustration of a conventional LLM answer (left) and an answer by a reasoning model (right).
提高(推理)任务性能本质上只有两种方式:训练扩展和推理扩展。
There are essentially two ways to improve (reasoning) task performance: training scaling and inference scaling.
图 4:训练扩展和推理扩展是提高 LLM 和推理模型问题解决能力的两种方式。图基于《Learning to reason with LLMs》绘制。
Figure 4: Training and inference-scaling are two ways to improve LLM and reasoning model problem-solving capabilities. Plot based on Learning to reason with LLMs
简而言之,DeepSeek-R1 提出使用带可验证奖励的强化学习(RLVR)来训练大语言模型(LLM),使其成为推理模型。RLVR 是一种为可验证数据领域提供奖励信号(0=不正确,1=正确)的技术。这里的可验证数据领域包括数学(我们可以使用符号数学检查器如 SymPy 或 WolframAlpha 来检查结果)和代码(我们可以使用编译器或单元测试,或集成平台如 LeetCode)来检查正确性。
In a nutshell, DeepSeek-R1 proposed training an LLM using reinforcement learning with verifiable rewards (RLVR) to turn it into a reasoning model. RLVR is a technique to provide a reward signal (0=incorrect and 1=correct) for verifiable data domains. These verifiable data domains here are math (we can use a symbolic math checker like SymPy or WolframAlpha to check results) and code (we can use a compiler or unit tests, or integrated platforms like LeetCode) to check for correctness.
图 5:RLVR 训练过程中准确率和格式奖励的示意图。
Figure 5: Illustration of accuracy and format rewards during RLVR training.
值得注意的是,推理轨迹本身并未用于训练或更新模型。尽管他们尝试使用这种中间响应信息进行训练,但 DeepSeek-R1 论文报告称这对模型训练没有帮助,因此最终未使用。(是否以及如何通过过程奖励模型将中间推理轨迹纳入训练信号,是一个活跃的研究领域。)
Notably, the reasoning trace itself was not used for training or updating the model. Although they tried to use this intermediate response information for training, the DeepSeek-R1 paper reported that it wasn’t helpful for the model training, so it was ultimately not used. (Whether and how to incorporate intermediate reasoning traces in the training signal via process reward models is an active area of research.)
图 6:在 RLVR 过程中,中间推理轨迹被忽略;只有最终答案和响应格式决定奖励。
Figure 6: The intermediate reasoning trace is ignored during RLVR; only the final answer and response format determine the reward.
无论如何,如图 7 所示,仅基于输出奖励进行训练,就足以让模型学会如何推理问题,即学会编写中间解释、回溯和自我纠正。模型意识到自己犯错并自我纠正的这些时刻被称为“顿悟”时刻。
Anyway, just training on the output rewards alone, as Figure 7 shows, turned out to be sufficient for the model to learn how to reason through a problem, meaning that it would learn to write intermediate explanations, backtrack, and self-correct itself. These moments when the model realizes that it made a mistake and self-corrects itself are called “Aha” moments.
图 7:一个“顿悟”时刻的示例,推理模型在生成最终答案之前注意到其中间推理中的错误并加以纠正。
Figure 7: An example of an aha moment, where a reasoning model notices an error in its intermediate reasoning and corrects it before producing the final answer.
顺便提一下,虽然 DeepSeek-R1 无疑更受欢迎,并且是引发对可验证奖励强化学习和推理模型开发兴趣的论文,但还有另一篇论文 Kimi K1.5,于同一天(2025 年 1 月 22 日)发表在 arXiv 上。此外,RLVR 这一术语早在两个月前就在 Tülu 3: Pushing Frontiers in Open Language Model Post-Training 中被提出。
By the way, while DeepSeek-R1 is inarguably the more popular paper, and the paper that created excitement around reinforcement learning with verifiable rewards and the development of reasoning models, there is another paper, Kimi K1.5, published on exactly the same day on arXiv (22 Jan 2025). Also, the term RLVR was already coined two months earlier in Tülu 3: Pushing Frontiers in Open Language Model Post-Training.
DeepSeek R1 最终更受欢迎的一个原因是,它证明了仅通过纯强化学习(RL)就能实现推理行为。
One reason why the DeepSeek R1 is ultimately the more popular paper is that it demonstrated that reasoning behavior can be achieved with pure reinforcement learning (RL).
图 8:DeepSeek-R1-Zero 直接对预训练基础模型应用 RLVR,无需监督微调。
Figure 8: DeepSeek-R1-Zero applies RLVR directly to the pretrained base model without supervised fine-tuning.
例如,Tülu 3 和 Kimi K1.5 在监督微调(SFT)模型之上应用了强化学习。DeepSeek-R1 模型也是从 DeepSeek-V3 基础模型的 SFT 检查点训练的,并且包含一个使用纯 RLVR 训练的 DeepSeek-R1-Zero 变体。R1 Zero 比 R1 弱,但它表明 RLVR 足以教会模型生成和使用推理轨迹。
For instance, Tülu 3 and Kimi K1.5 applied reinforcement learning on top of a supervised fine-tuned (SFT) model. The DeepSeek-R1 model was also trained from an SFT checkpoint of the DeepSeek-V3 base model, and it included a DeepSeek-R1-Zero variant trained with pure RLVR. R1 Zero is a weaker model than R1, but it showed that RLVR is sufficient for teaching the model to generate and use reasoning traces.
虽然 R1-Zero 更像是一个概念验证模型,但请注意,完整的 DeepSeek-R1 推理模型训练流程通常是多阶段的,并且如上所述,稍微复杂一些。
While R1-Zero was more of a proof-of-concept model, note that the full DeepSeek-R1 reasoning model training pipeline is usually multi-stage and a bit more complicated, as mentioned above.
图 9:更详细的推理模型训练流程。此图展示了各种 DeepSeek-R1 模型。更多细节,请参阅我的另一篇文章:《理解推理大语言模型》。
Figure 9: More detailed reasoning model training pipeline. This one depicts the various DeepSeek-R1 models. For more details, see my other article: Understanding Reasoning LLMs.
顺便说一句,如今大多数大语言模型实际上都是推理模型,这意味着它们采用了与 DeepSeek-R1 类似的方式,通过某种形式的基于强化学习的验证(RLVR)进行训练。
By the way, most of today’s LLMs are effectively reasoning models, meaning they have been trained in a similar fashion to DeepSeek-R1 using a form of RLVR.
除了通过训练改善推理行为外,提升模型性能的另一个杠杆是推理时的算力扩展(inference compute scaling)。简而言之,这意味着我们在训练模型之后、使用过程中投入更多算力,以获得更好的答案。
Next to improving reasoning behavior through training, another lever for improving model performance is inference compute scaling. In short, this means that we are spending more compute after training the model, during usage, to get better answers.
这本身就是一个完整的主题,你可以阅读我的《LLM 推理模型的状态》一文,获取更详细的介绍:
This is a whole topic by itself, and you could read through my The State of LLM Reasoning Model Inference for a more detailed rundown:
#### 面向 LLM 推理的强化学习现状 [Sebastian Raschka 博士 · 2025 年 4 月 19 日 阅读全文](https://magazine.sebastianraschka.com/p/the-state-of-llm-reasoning-model-training)
#### The State of Reinforcement Learning for LLM Reasoning [Sebastian Raschka, PhD · April 19, 2025 Read full story](https://magazine.sebastianraschka.com/p/the-state-of-llm-reasoning-model-training)
下面我将尝试总结作为背景信息最需要提及的内容。
I will try to summarize what’s most essential to mention as background info below.
首先,使用 RLVR 训练模型已经隐式地导致了某种形式的推理时规模扩展,因为推理模型在推理时通常比传统 LLM 输出更多的词元,这意味着我们在推理时投入了更多的算力。
First, training a model with RLVR is already implicitly leading to a form of inference scaling, since reasoning models usually output more tokens during inference compared to conventional LLMs, and that means we are spending more compute during inference.
其次,我们可以通过推理努力程度进一步调整输出长度,但稍后会详细讨论。
Second, we can further adjust this output length via reasoning effort levels, but more on that later.
第三,还有许多其他的推理缩放技术。其中一种流行的方法是自洽性(self-consistency),它通常以多数投票的形式实现:模型被多次查询,最终答案通过多数投票选出。
Third, there are many additional inference scaling techniques. A popular one is self-consistency, which is often implemented as a form of majority voting where the model is queried multiple times, and the final answer is selected via majority vote.
图 10:自洽性示例,一种流行的推理缩放技术。
Figure 10: An example of self-consistency, a popular inference scaling technique.
这种方法既适用于传统 LLM,也适用于推理模型。此外,该方法可以按需使用,并且可以与推理训练结合使用。一个很好的例子是 DeepSeekMath-V2,研究人员在推理模型(专门针对数学)之上应用了极端的推理缩放,在具有挑战性的数学奥林匹克类问题上取得了最先进的性能。
This can be applied to conventional LLMs as well as reasoning models. Also, this method can be used on demand and in addition to reasoning training. A good example of that is DeepSeekMath-V2, where the researchers applied extreme inference-scaling on top of a reasoning model (specialized for math) to achieve state-of-the-art performance on challenging math olympiad-type problems.
图 11:两种推理扩展(自洽性和自我精炼)结合使用以提高数学性能。图改编自 DeepSeekMath-V2:迈向自验证的数学推理
Figure 11: Two types of inference scaling (self-consistency and self-refinement) used together to improve math performance. Figure adapted from DeepSeekMath-V2: Towards Self-Verifiable Mathematical Reasoning
但同样,我将参考我的另一篇文章《LLM 推理模型推理状态》以了解其他技术的概述:
But again, I will refer to my other article, The State of LLM Reasoning Model Inference for an overview of other techniques:
#### LLM 推理模型推理状态 [Sebastian Raschka, PhD · 2025 年 3 月 8 日 阅读全文](https://magazine.sebastianraschka.com/p/state-of-llm-reasoning-and-inference-scaling)
#### The State of LLM Reasoning Model Inference [Sebastian Raschka, PhD · March 8, 2025 Read full story](https://magazine.sebastianraschka.com/p/state-of-llm-reasoning-and-inference-scaling)
你可能在之前的“顿悟时刻”图中见过 <think></think> 令牌。我也在下面附上了相应的图,这样你就不必一直向上滚动。
You may have seen the <think></think> tokens in the earlier “Aha moments” figure. I also included the corresponding figure below so you don’t have to scroll all the way up.
图 12:推理模型中的常见格式令牌。
Figure 12: Common formatting tokens in reasoning models.
这些 <think> 和 </think> 标签在推理能力方面是装饰性的。它们不会让模型进行推理,也不是获得良好推理性能所必需的。人们可以在没有这些分隔符的情况下训练相同的模型,并且很可能达到类似的基准性能。
These <think> and </think> tags are cosmetic with respect to reasoning ability. They do not make the model reason, and they are not required to achieve good reasoning performance. One could train the same model without these delimiters and likely reach similar benchmark performance.
这些 <think> 标签或令牌的主要目的是标记推理轨迹的开始和结束位置,以便训练流程或用户界面将其与最终答案分开,并可选地对用户隐藏。(像 ChatGPT 或 Codex 这样的用户界面通常会这样做。)
The purpose of these <think> tags or tokens is mainly to mark where the reasoning trace begins and ends so that the training pipeline or user interface can separate it from the final answer and optionally hide it from the user. (UIs like ChatGPT or Codex usually do this.)
这里的要点是,<think> 令牌并没有赋予模型“思考”或推理或更好推理的能力。人们可以在没有这些 <think> 令牌的情况下训练相同的模型,并达到类似的基准性能。
The point here is that the <think> tokens are not giving the model the ability to “think” or reason or reason better. One could train the same models without such <think> tokens and reach similar benchmark performance.
此外,字面字符串 <think> 和 </think> 也没有任何特殊之处。另一对分隔符也可以起到同样的作用。
There is also nothing special about the literal strings <think> and </think>. Another pair of delimiters could serve the same purpose.
顺便提一下,这种实现方式通常是在 RLVR 阶段添加一个格式奖励。也就是说,不仅根据答案正确性奖励模型,还会为使用 <think> 令牌提供额外奖励,从而鼓励模型使用这些令牌。
By the way, the way this is implemented is typically by adding a formatting reward during the RLVR stage. So instead of just rewarding the model based on answer correctness, one would provide additional reward for the use of <think> tokens, which in turn encourages the model to use those.
例如,在 DeepSeek-R1 中,总奖励的计算方式为
In DeepSeek-R1, for example, the overall reward was calculated as
其中格式奖励是一个简单的基于规则的检查,鼓励模型将推理过程放在其中:
where the format reward was a simple rule-based check that encouraged the model to place its reasoning inside:
第一代推理模型是专用推理模型。我的意思是,当时有一个 DeepSeek-V3 基础模型和一个独立的 DeepSeek-R1 推理模型。
The first generation of reasoning models was dedicated reasoning models. With that, I mean that there was a DeepSeek-V3 base model and a separate DeepSeek-R1 reasoning model.
无论提示词是什么,R1 通常都会输出非常冗长的回答,消耗大量词元,即使对于简单的提示词也是如此。它还缺乏内置的关闭推理模式的选项。
No matter what the prompt is, R1 generally outputs very verbose responses using lots of tokens, even for simple prompts. It also lacks a built-in option to turn off the reasoning mode.
图 13:推理模型非常冗长,即使对于最简单的提示词也是如此。
Figure 13: Reasoning models are very verbose, even for the simplest prompts.
后来的模型,如 Qwen3 等,尝试了混合方法,同一个模型可以根据需求表现得像常规的指令微调模型或推理模型。
Later models, like Qwen3 and others, experimented with hybrid approaches, where the same model can behave like a regular instruction fine-tuned model or a reasoning model on demand.
The first generation of reasoning models was dedicated reasoning models. With that, I mean that there was a DeepSeek-V3 base model and a separate DeepSeek-R1 reasoning model.
在 Qwen3 中,这是通过分词器(tokenizer)使用 enable_thinking=True 或 enable_thinking=False 来处理的。在底层,设置 enable_thinking=False 实际上是在助手响应的开头添加一个空的 <think></think> 部分,以关闭 Qwen3 的推理("思考")模式。
In Qwen3, this is handled via the tokenizer using enable_thinking=True or enable_thinking=False. Under the hood, setting enable_thinking=False essentially adds an empty <think></think> section to the beginning of the assistant response to turn off Qwen3's reasoning ("thinking") mode.
图 14:Qwen3 0.6B 推理模型在 thinking=False 和 thinking=True 时的响应。(左侧界面中空的 <think></think> 标签被隐藏,因为它们是修改后的输入提示的一部分,而非生成的答案。)
Figure 14: Response of Qwen3 0.6B reasoning model with thinking=False and thinking=True. (The empty <think></think> tags are hidden in the interface on the left as they are part of the modified input prompt, not the generated answer.)
在训练过程中是如何实现这一点的,使得模型在推理时支持这种切换,如上图所示?
How is this implemented during training, such that the model supports this toggle during inference time, as shown in the figure above?
简而言之,正如 Qwen3 技术报告中所解释的,这种开/关行为主要通过监督微调(SFT)引入,并在其最大的旗舰模型中通过通用强化学习(RL)加以强化。
In short, as explained in the Qwen3 technical report, this on/off behavior is introduced primarily through supervised fine-tuning (SFT) and then reinforced during general RL in their largest flagship models.
例如,在初始推理模型通过长思维链 SFT 和推理强化学习训练之后,他们增加了一个“思考模式融合”阶段。在这个额外的 SFT 阶段,模型同时看到思考和非思考的示例:
For instance, after the initial reasoning model is trained via long-chain-of-thought SFT and reasoning RL, they add a “Thinking Mode Fusion” stage. During this additional SFT stage, the model sees both thinking and non-thinking examples:
* /think: <think>{reasoning}</think>{answer}
* /think: <think>{reasoning}</think>{answer}
思考是默认行为,因此 /think 也可以省略。后续的通用强化学习阶段进一步强化了这种模式和格式遵循。
Thinking is the default behavior, so /think can also be omitted. The subsequent general RL stage further reinforces this mode and format following.
这些 /think 和 /no_think 标志是一个“软”开关。然而,前面提到的 enable_thinking=False 设置,在 False 情况下强制添加空的 <think></think>,则充当“硬”开关。
These /think and /no_think flags are a “soft” switch. However, the enable_thinking=False setting mentioned earlier, which force-adds the empty <think></think> in the False case, acts then as a “hard” switch.
图 15:Qwen3 训练流程中的“思维模式融合”,用于实现推理模式的开关切换。
Figure 15: “Thinking Mode Fusion” in Qwen3’s training pipeline to enable the reasoning mode on and off switch.
换句话说,分词器不会在查询中添加 /no_think。它直接在助手响应的开头填充空的 <think></think> 部分。模型只看到生成的词元,并直接继续回答。
In other words, the tokenizer does not add /no_think to the query. It directly fills in the empty <think></think> section at the beginning of the assistant response. The model only sees the resulting tokens and continues directly with the answer.
总之,这种开关切换本质上是 GPT-5.6 等模型中推理努力等级的简化版本,我将在下一节中介绍。
Anyway, this on-and-off toggle is essentially a simplified version of the reasoning effort levels in GPT-5.6 and others, which I cover in the next section.
在本节中,我想简要概述不同的推理努力开关可能如何实现,这些开关已在 GPT-5 等模型中引入,并且如今几乎出现在所有旗舰模型中。
In this section, I want to provide a brief overview of how the different reasoning effort toggles may be implemented, which have been introduced in models like GPT-5 and are present in pretty much any flagship model today.
具体来说,在本文开头,我展示了一张来自 Codex GPT-5.6 界面的图,该界面允许用户选择多个推理“努力”设置。
Concretely, at the beginning of this article, I showed a figure from the Codex GPT-5.6 interface that lets users select multiple reasoning “effort” settings.
图 16:GPT-5.6 提供了六种推理努力设置,范围从 Light 到 Ultra。
Figure 16: GPT-5.6 exposes six reasoning effort settings, ranging from Light to Ultra.
接下来的小节将说明这些设置可能如何实现。然后,在下一节中,我将介绍与该主题相关的一些更有趣的研究论文。
The following subsection will illustrate how these settings may be implemented. Then, in the next section, I will go over some of the more interesting research papers related to this topic.
不幸的是,OpenAI 并未公开其努力设置的具体实现细节,但有一些证据可以用来进行有根据的猜测。
Unfortunately, the implementation details of their effort settings are not shared by OpenAI, but there is some evidence out there that can be used for educated guesses.
例如,通过他们去年发布的开源 gpt-oss 模型(我在《从 GPT-2 到 gpt-oss:分析架构进展》中写过),我们知道 OpenAI 允许我们通过系统提示(“推理努力:低/中/高”)来切换推理努力设置,该提示会前置到每个提示中。
For instance, via their open-source gpt-oss models from last year (I wrote about them in From GPT-2 to gpt-oss: Analyzing the Architectural Advances), we know that OpenAI allows us to toggle the reasoning effort setting via the system prompt (”Reasoning effort: low/medium/high”) that is prepended to each prompt.
图 17:gpt-oss 聊天模板在将提示发送给同一模型之前,会将选定的推理努力插入到系统消息中。
Figure 17: The gpt-oss chat template inserts the selected reasoning effort into the system message before sending the prompt to the same model.
正如预期的那样,推理努力直接影响响应长度和准确性,如下所示。
As expected, the reasoning effort directly affects the response length and accuracy, as shown below.
图 18:不同推理努力下 gpt-oss 模型的响应长度和质量(来自模型卡的注释图)
Figure 18: Response length and quality of gpt-oss models under different reasoning efforts (annotated figure from the model card)
据推测,他们的 GPT 5 模型,包括最近的 GPT 5.6 模型,采用了类似的方法。
Presumably, their GPT 5 models, including the recent GPT 5.6 models, use a similar approach.
顺便提一下,注意上图中不同努力设置如何缩放响应长度。努力水平似乎与 token 使用量直接相关,而 token 使用量又与准确性相关。有可能提出超出“高”水平的努力设置,但我认为性能会在某个点饱和。这种饱和在 GPT 5.6 Sol 模型中更为明显,这也表明增加推理预算在某个点可能变得不经济。
By the way, note how different effort settings scale the response length in the figure above. The effort level seems directly correlated to token usage, which in turn seems correlated to accuracy. It might be possible to come up with effort settings beyond the “high” one, but I assume performance would saturate at some point. This saturation can be seen more clearly for the GPT 5.6 Sol model, which also shows that increasing reasoning budgets can become uneconomical at some point.
图 19:推理努力增加了 API 成本和编码智能体性能,但在最高的 GPT-5.6 设置下收益递减。图基于 Artificial Analysis 编码智能体指数 v1.1。
Figure 19: Reasoning effort increases both API cost and coding-agent performance, with diminishing returns at the highest GPT-5.6 settings. Figure based on the Artificial Analysis Coding Agent Index v1.1.
另一个很好的、非常新的数据点展示了推理努力、token 使用和基准性能之间的关系,即 Thinking Machine Labs 本周发布的新开放权重模型 Inkling。
Another good, very recent data point that shows the relationship between reasoning effort, token usage, and benchmark performance is this week’s new open-weight Inkling release by Thinking Machine Labs.
图 20:增加 Inkling 的努力水平通常会增加生成的 token 和基准性能,但在更高的努力水平下收益递减或不均匀。图来自 Inkling 发布博客。
Figure 20: Increasing the Inkling effort level generally increases generated tokens and benchmark performance, with diminishing or uneven gains at higher effort. Figure from the Inkling announcement blog.
如本节所讨论的,在推理过程中,推理努力水平可以简单地通过系统提示来控制。(ChatGPT 界面可能只是将菜单选择映射到系统提示。)然而,这对于任意模型并不适用,需要对训练流程进行某些修改,这将在接下来讨论。
As discussed in this section, during inference, the reasoning effort level can simply be controlled via a system prompt. (The ChatGPT UI presumably simply maps the menu choice to a system prompt.) However, this would not work for an arbitrary model and requires certain modifications to the training pipeline, which will be discussed next.
尽管训练细节并未公开,无论是 GPT-5.6 还是开源的 gpt-oss 模型,通常在后训练期间,推理努力标签会包含在提示中。
While the training details are not public, neither for GPT-5.6 nor the open-source gpt-oss models, typically, the reasoning effort label is included in prompts during post-training.
通常有两种实现方式。
There are typically two ways to implement this.
第一种,我们可以将其作为 RLVR 过程的一部分,并在使用不同系统提示时应用不同的长度惩罚。例如,当提示为“推理努力:低”时使用高长度惩罚,而当提示为“推理努力:高”时使用轻微或无惩罚。
First, we can implement it as part of the RLVR process and apply a different length penalty when different system prompts are used. For example, a high length penalty when “Reasoning effort: low” and a mild or no penalty when “Reasoning effort: high”.
第二种,我们可以在 RLVR 之后通过监督微调(SFT)对模型进行微调,使其遵循不同的努力指令。
Second, we can fine-tune the model after RLVR to follow different effort instructions via supervised fine-tuning (SFT).
例如,在核心 RLVR 阶段之后,在 SFT 期间,训练数据集中的提示与展示所需推理量的目标响应配对。(目标响应可能由人类编写、由另一个模型生成,或生成后经过筛选。)
For instance, after the core RLVR stage and during SFT, the prompts in the training dataset are paired with target responses that exhibit the desired amount of reasoning. (The targets may be written by humans, generated by another model, or generated and then filtered.)
图 21:努力条件化的 RLVR 和 SFT 示意图。(这是一种可能的实现方式,并非对 OpenAI 训练流程的确认描述。)
Figure 21: Illustration of effort-conditioned RLVR and SFT. (This is a possible implementation, not a confirmed description of OpenAI’s training pipeline.)
在 SFT 阶段,模型直接从训练样本中学习努力标签与目标推理长度之间的关联。而基于 RL 的实现则会将努力标签和预算感知奖励置于 RLVR 阶段。这两种方法也可以结合使用,我怀疑 gpt-oss 和 GPT 5.6 都采用了这种结合方式(注意,GPT 5.6 中的努力设置可能只是针对给定用户查询更改系统提示)。
During this SFT stage, the model learns the association between the effort label and the target reasoning length directly from the training examples. An RL-based implementation would instead place the effort labels and budget-aware reward inside the RLVR stage. The two approaches could also be combined, which I suspect was done for both gpt-oss and GPT 5.6 (note that effort settings in GPT 5.6 are likely just changing the system prompt for a given user query).
刚刚发布的 Inkling 技术报告给出了一个规模不大但较为具体的努力水平训练示例。
The just-released Inkling technical report gives a small but somewhat concrete example of effort-level training.
图 22:Inkling 在 0.2 到 0.99 之间扫描连续的努力值;更高的努力通常会产生更长的响应和更高的基准分数。
Figure 22: Inkling sweeps a continuous effort value between 0.2 and 0.99; higher effort generally produces longer responses and higher benchmark scores.
在大规模强化学习过程中,他们对每个样本做了两件事:
During large-scale RL, they did two things for each sample:
1. 在系统消息中指定所需的努力水平。
1. Specified the desired effort level in the system message.
The just-released Inkling technical report gives a small but somewhat concrete example of effort-level training.
2. 调整分配给每个生成 token 的成本。
2. Adjusted the cost assigned to each generated token.
从概念上讲,奖励可能如下所示:
Conceptually, the reward likely looked something like this:
\[ R(e) = R_{\text{task}} - \lambda(e) N_{\text{tokens}} \]
\[ R(e) = R_{\text{task}} - \lambda(e) N_{\text{tokens}} \]
这里,_e_ 是请求的努力水平,λ(e) 控制 token 惩罚。
Here, _e_ is the requested effort level and λ(e) controls the token penalty.
* 低努力使用更大的每 token 成本,鼓励更短的推理轨迹。
* Low effort uses a larger per-token cost, encouraging shorter reasoning traces.
* 高努力使用更小的每 token 成本,允许模型花费更多 token。
* High effort uses a smaller per-token cost, allowing the model to spend more tokens.
然后,在推理时,Inkling 会收到一条系统消息,如“思考努力水平:0.8”,并相应地调整其 token 使用量。Inkling 与 gpt-oss 和 GPT-5.6 等模型之间的区别在于,努力标签是 0 到 1 之间的连续数字,而不是诸如低、中、高之类的序数标签。
Then, at inference time, Inkling receives a system message such as "Thinking effort level: 0.8", and adjusts its token usage accordingly. The difference between Inkling and models such as gpt-oss and GPT-5.6 is that the effort label is a continuous number between 0 and 1 instead of ordinal labels such as low, medium, and high.
这使得 Inkling 的努力水平条件化主要发生在推理强化学习阶段,而不仅仅是在后期的 SFT 阶段。
This places Inkling's effort-level conditioning primarily in the Reasoning RL stage, not only in the later SFT stage.
不过,他们没有透露确切的奖励公式、token 成本系数,也没有透露努力条件化是否也包含在 SFT 中。
They do not disclose the exact reward formula, token-cost coefficients, or whether effort conditioning was also included in SFT, though.
在继续讨论推理努力(reasoning effort)相关论文之前,我想简要地将本节与前面的“2.3 推理扩展简述”一节联系起来。
Before moving on to the reasoning effort papers, I want to briefly connect this section back to the earlier “2.3 Inference scaling in a nutshell” section.
之前,我将扩展分为训练算力扩展和推理时扩展。GPT-5.6 的界面提供了一个很好的方式来展示二者的区别,如下所示。
Earlier, I separated scaling into training compute scaling and inference-time scaling. The GPT-5.6 interface provides a nice way to illustrate the difference, as shown below.
在左侧,选择 Luna、Terra 或 Sol 会改变模型本身。粗略类比,这对应于训练算力扩展。这些是独立训练的模型。在固定的训练方案和数据集大小下,更大的模型需要更多的训练算力,通常每个生成 token 也需要更多算力。
On the left, selecting Luna, Terra, or Sol changes the model itself. As a rough analogy, this corresponds to training compute scaling. These are separate trained models. At a fixed training recipe and dataset size, a larger model requires more training compute. It also generally requires more compute per generated token.
在右侧,我们保持模型固定,仅改变推理努力。这就是推理时扩展。模型权重保持不变,但允许模型在解答时花费更少或更多的 token。
On the right, we keep the model fixed and only change the reasoning effort. This is inference-time scaling. The model weights stay the same, but the model is allowed to spend fewer or more tokens working on the answer.
图 23:模型选择菜单和推理努力菜单对应两个不同的扩展轴。选择 Luna、Terra 或 Sol 会改变模型,而改变推理努力则调整固定模型的推理时算力。
Figure 23: The model selection and reasoning effort menus correspond to two different scaling axes. Selecting Luna, Terra, or Sol changes the model, whereas changing the reasoning effort adjusts the inference-time compute for a fixed model.
一个小的术语注意事项是,从菜单中选择不同的模型并不是当时的训练扩展。训练已经发生了。更好的理解是,模型菜单是在选择不同训练规模下产生的模型。
One small terminology caveat is that selecting a different model from the menu is not training scaling at that moment. The training has already happened. It is better to think of the model menu as selecting among models that were produced at different training scales.
下面的 Artificial Analysis 结果展示了这两个轴在实际中的交互。
The Artificial Analysis results below show how these two axes interact in practice.
每条蓝色曲线对应一个模型:Luna、Terra 或 Sol。通过增加推理努力沿曲线移动是推理扩展。从一条模型曲线移动到另一条对应模型扩展,我在这里将其作为训练扩展的实际代理。
Each blue curve corresponds to one model, Luna, Terra, or Sol. Moving along a curve by increasing the reasoning effort is inference scaling. Moving from one model curve to another corresponds to model scaling, which I use here as a practical proxy for training scaling.
正如预期,两种方法都能提高基准分数,但也增加了成本。更有趣的是,曲线有重叠。例如,较小的模型在较高的推理努力下有时可以达到与较大模型在较低推理努力下相似的分数。
As expected, both approaches can improve the benchmark score, but they also increase the cost. More interestingly, the curves overlap. For instance, a smaller model at a higher reasoning effort can sometimes reach a similar score as a larger model at a lower reasoning effort.
图 24:GPT-5.6 模型系列在 Artificial Analysis 编码智能体指数上的训练扩展与推理扩展。沿每条模型曲线移动对应增加推理努力。在 Luna、Terra 和 Sol 曲线之间移动对应选择不同的模型。
Figure 24: Training scaling and inference scaling for the GPT-5.6 model family on the Artificial Analysis Coding Agent Index. Moving along each model curve corresponds to increasing the reasoning effort. Moving across the Luna, Terra, and Sol curves corresponds to selecting a different model.
顺便提一下,该图中的 x 轴显示的是 API 成本而非原始算力。API 成本是一个实用的实际度量,但它也取决于提供商的定价和生成的 token 数量。此外,这些曲线的确切形状是特定于基准的。
By the way, the x-axis in this figure shows API cost rather than raw compute. The API cost is a useful practical measure, but it also depends on the provider’s pricing and the number of generated tokens. Also, the exact shape of these curves is benchmark-specific.
因此,模型大小和推理努力构成两个独立的旋钮。我们可以使用更大的模型、增加推理努力,或两者结合。哪种组合最佳取决于所需的准确性、成本和延迟。
So, the model size and reasoning effort form two separate knobs. We can use a larger model, increase the reasoning effort, or combine both. Which combination is best depends on the desired accuracy, cost, and latency.
到目前为止,本文应该让您对推理努力模式的工作原理及其实现方式有了相当扎实的理解。如果您时间有限,这是一个很好的收尾点。否则,如果您想深入了解一些近期开放权重模型的细节,请继续阅读!
So far, the article should give you a pretty solid understanding of how reasoning effort modes work and how they are implemented. This is a fine point to wrap the article if you are short on time. Otherwise, if you want to look into some of the nitty-gritty details of some of the recent open-weight models, please read on!
[除非你对一些额外细节感兴趣,否则可以跳过本节]
[This section is fine to skip unless you are interested in some additional details]
第 5 节描述了训练推理努力控制的两种可能方式,即基于努力条件的监督微调和具有不同 token 成本的强化学习。最初,我想涵盖关于实现推理预算的替代方法的研究文章。然而,在阅读了这些文章中的大部分之后,它们似乎更像是概念验证,在实践中可能有效也可能无效。
Section 5 described two possible ways to train reasoning-effort controls, namely effort-conditioned supervised fine-tuning and reinforcement learning with different token costs. Originally, I wanted to cover research articles on alternative ways to implement reasoning budgets. However, reading through most of these articles, they seemed more like proofs-of-concept that may or may not work well in practice.
因此,我没有涵盖这些内容,而是决定稍微转向,涵盖最先进且著名的开源权重(旗舰)LLM 所使用的那些配方。对于这些模型,至少有证据表明这些方法在实践中有效。
So, instead of covering those, I decided to pivot a bit and cover those recipes used by state-of-the-art and notable open-weight (flagship) LLMs. For these models, there is at least evidence that the methods work in practice.
这留下了六个例子:DeepSeek V4、Nemotron 3 Ultra、Kimi K2.5、GLM-5、Qwen3 和 Inkling。它们的报告详细程度不同,但每个都贡献了一个有用的变体。(我排除了那些仅在用户界面中显示努力设置而未解释该行为如何训练的模型。)
This leaves six examples. DeepSeek V4, Nemotron 3 Ultra, Kimi K2.5, GLM-5, Qwen3, and Inkling. They have different levels of detail in their reporting, but each contributes a useful variation. (I exclude models whose reports only show an effort setting in the user interface without explaining how that behavior was trained.)
让我们从 DeepSeek V4 技术报告开始,该报告描述了三种模式的使用:
Let’s start with the DeepSeek V4 technical report, which describes the use of three modes:
* **非思考模式**:直接生成响应,不包含推理过程。 * **高思考模式**:经典方法,模型将推理过程放在 <think> 和 </think> 标签之间。这与本文开头 DeepSeek R1 部分(第 2 节)讨论的内容类似。 * **最大思考模式**:与上述相同,但添加了特殊的系统指令。(下文将详细说明。)
* Non-think produces a direct response without a reasoning trace. * Think High is the classic approach where the model places the reasoning trace between <think> and </think> tags. This is similar to what was discussed in the DeepSeek R1 section (section 2) at the beginning of this article. * Think Max is the same as above but adds a special system instruction. (More on that below.)
最大思考模式的额外系统提示指令以“推理努力:绝对最大,不允许捷径”开头。
The additional system prompt instruction for Think Max starts with “Reasoning Effort: Absolute maximum with no shortcuts permitted.”
Let’s start with the[](https://www.google.com/url?q=https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro/resolve/main/DeepSeek_V4.pdf?download%3Dtrue&sa=D&source=editors&ust=1784335866750173&usg=AOvVaw1sM7dNkExL5WhfZEIXpAvQ)DeepSeek V4 technical report, which describes the use of three modes:
* Non-think produces a direct response without a reasoning trace.
图 25:DeepSeek V4 文档中的推理努力控制概览
Figure 25: Reasoning effort control overview from the DeepSeek V4 documentation
乍一听,这似乎只是一个简单的提示工程技巧,但实际上,这个提示背后有着不同的训练设置。也就是说,每种模式都有各自的上下文窗口和长度惩罚(遗憾的是,该报告未详细说明长度惩罚的具体实现)。与 Think High 相比,Think Max 拥有更长的上下文窗口和更小的长度惩罚,这为其继续推理提供了更大的空间。
At first, this sounds like a simple prompt engineering trick, but this prompt is actually backed by a different training setup. That is, each mode uses its own context window and length penalty (unfortunately, the exact length penalty implementation is not detailed in this report). Think Max receives a longer context window and a smaller length penalty than Think High, which gives it more room to continue reasoning.
因此,系统指令选择的是后训练期间创建的行为。将相同的指令添加到任意模型上,不会产生相同的效果。
So, the system instruction selects a behavior that was created during post-training. Adding the same instruction to an arbitrary model would not have the same effect.
图 26:DeepSeek V4 在报告的不同部分分别描述了三种努力模式以及更大的教师池。教师池包含十多个领域专家。报告未披露这些教师如何映射到 Non-think、Think High 和 Think Max。
Figure 26: DeepSeek V4 describes the three effort modes and the larger teacher pool in separate parts of the report. The teacher pool contains more than ten domain specialists. The report does not disclose how these teachers map to Non-think, Think High, and Think Max.
遗憾的是,公开且非常详细的 DeepSeek V4 报告并未将推理模式和领域专家的描述以足够的细节联系起来,无法重建确切的教师分配。
Unfortunately, the public, and otherwise very detailed DeepSeek V4 report does not connect the descriptions of the reasoning mode and domain specialists in enough detail to reconstruct the exact teacher assignment.
然而,报告指出,支持不同推理努力水平的最终模型是通过从上述教师进行在线策略蒸馏(on-policy distillation)创建的。
However, the report states that the final model, which supports different reasoning effort levels, was created via on-policy distillation from said teachers.
总结来说,DeepSeek V4 在后训练期间开发了三种推理专家。从基础模型开始,应用监督微调,随后通过 GRPO 进行 RLVR。每种模式的 RL 配置不同。特别是,每个专家使用自己的上下文窗口和长度惩罚,而 Think Max 还额外接收特殊的系统指令。
To summarize, DeepSeek V4 develops the three reasoning specialists during post-training. Starting from the base model, it applies supervised fine-tuning followed by RLVR via GRPO. The RL configuration differs for each mode. In particular, each specialist uses its own context window and length penalty, while Think Max additionally receives a special system instruction.
然后,包括领域专家在内,不同的推理模式专家被蒸馏到一个支持所有三种努力模式的单一检查点中。
Then, including domain specialists, the different reasoning mode specialists are distilled into a single checkpoint that supports all three effort modes.
Nemotron 3 Ultra 技术报告描述了三种设置:reasoning-off、regular 和 medium-effort,与上一节中的 DeepSeek V4 类似。与 regular 相比,medium-effort 是更便宜的推理模式。NVIDIA 在 SFT 阶段使用 GPT-OSS-120B 在其 medium-effort 模式下生成的示例引入该模式,并在 RLVR 期间进一步优化。大约 2.5% 的 RLVR 提示使用 medium-effort(这对应于对其奖励进行的基于长度的调整)。
The Nemotron 3 Ultra technical report describes three settings called reasoning-off, regular, and medium-effort, analogous to DeepSeek V4 in the previous section. Medium-effort is the cheaper reasoning mode compared with regular. NVIDIA introduces this mode during SFT using examples generated by GPT-OSS-120B in its medium-effort mode, and then further optimizes it during RLVR. About 2.5% of the RLVR prompts use medium-effort (this corresponds to length-based adjustments applied to their rewards).
在推理时,所有三种模式均通过聊天模板进行选择。
At inference time, all three modes are selected through the chat template.
图 27:通过聊天模板设置 Nemotron 3 Ultra 的推理选项(示例来自官方模型卡)
Figure 27: Nemotron 3 Ultra reasoning settings via the chat template (examples from the official model card)
1) 常规模式为默认模式,使用 enable_thinking=True,这会使助手响应以 <think> 标签开头。
1) Regular is the default and uses enable_thinking=True, which starts the assistant response with an opening <think> tag.
2) 中等努力模式同时使用 enable_thinking=True 和 medium_effort=True,后者还会在最新用户消息后附加 {reasoning effort: efficient}。
2) Medium-effort uses enable_thinking=True together with medium_effort=True where the latter setting also appends {reasoning effort: efficient} to the latest user message.
顺便提一下,使情况更加复杂的是,常规模式和中等努力模式还可以与单独的推理时推理预算相结合。该预算充当外部停止机制。在发布的实现中,聊天客户端要求模型在接近所选令牌限制时结束推理轨迹。如果模型尚未输出 </think>,客户端将关闭推理块并继续生成以产生最终答案。学习到的努力模式决定模型如何使用其推理令牌,而预算则限制推理轨迹可以持续的时间。这使得可以根据期望的成本和准确性,将任一模式与更紧或更松的预算配对。
By the way, to further complicate things, the regular and medium-effort modes can also be combined with a separate inference-time reasoning budget. This budget acts as an external stopping mechanism. In the released implementation, the chat client asks the model to end the reasoning trace near the chosen token limit. If the model has not emitted </think>, the client closes the reasoning block and continues generation to produce the final answer. The learned effort mode determines how the model uses its reasoning tokens, while the budget constrains how long the reasoning trace can continue. This makes it possible to pair either mode with a tighter or looser budget depending on the desired cost and accuracy.
3) 关闭推理使用 enable_thinking=False,它会预填充一个空的 <think></think> 块(类似于第 4 节讨论的 Qwen3),以便模型直接进入最终响应。因此,这些是聊天模板控制,而非系统提示。
3) Reasoning-off uses enable_thinking=False, which prefills an empty <think></think> block (similar to Qwen3 discussed in section 4) so that the model proceeds directly to the final response. Thus, these are chat-template controls rather than system prompts.
上述推理控制由两个相关的 SFT 组件支持。第一个组件利用 GPT-OSS-120B 的轨迹引入中等努力行为,如前所述。第二个组件为困难推理预算做好准备。
The inference controls described above are backed by two related SFT components. The first introduces medium-effort behavior using GPT-OSS-120B traces, as discussed earlier. The second prepares the model for hard reasoning budgets.
为了构建这些训练数据,作者采用常规推理轨迹,在随机选择的 token 预算处截断,并保留原始最终答案。插入的 </think> token 在 SFT 损失中被掩蔽。因此,模型看到的示例是:在推理块被外部关闭后,它必须从不完整的推理轨迹过渡到答案。
To construct this training data, the authors take regular reasoning traces, truncate them at randomly selected token budgets, and keep the original final answers. The inserted </think> token is masked from the SFT loss. As a result, the model sees examples where it has to move from an incomplete reasoning trace to the answer after the reasoning block has been closed externally.
中等努力训练随后在 RLVR 期间继续进行。在数学、STEM 和编程任务中,约 2.5% 的 RL 提示使用中等努力设置。报告指出,该模式可以通过奖励超参数进行校准,其中基于长度的奖励调整提供了对成本-质量权衡的额外控制。
Medium-effort training then continues during RLVR. About 2.5% of the RL prompts use the medium-effort setting across math, STEM, and coding tasks. The report notes that the mode can be calibrated through reward hyperparameters where length-based reward adjustments provide additional control over the cost-quality trade-off.
图 28:Nemotron 3 Ultra 通过教师生成的 SFT 数据、随机预算截断以及 RLVR 期间的一小部分中等努力子集引入中等努力。
Figure 28: Nemotron 3 Ultra introduces medium effort with teacher-generated SFT data, random-budget truncation, and a small medium-effort subset during RLVR.
Kimi K2.5 技术报告讨论了一种名为“Token 高效强化学习”的训练方法,旨在降低推理开销。(尽管本周有 K3 的发布公告,但 K3 的推理开销方法并未公开披露,不过它可能与 K2.5 类似或相关。)
The Kimi K2.5 technical report discusses a training method called Token Efficient RL for lower reasoning effort. (While there was a K3 announcement this week, the reasoning-effort methodology of K3 is not publicly disclosed, but it could be similar or related to K2.5.)
报告提到,固定的 token 预算可能使推理模型过拟合于短解。这意味着模型变得更简洁(即更快、更便宜),但可能失去从额外推理时算力中获益的能力,从而表现不佳。
The report mentions that a fixed token budget can make a reasoning model overfit to short solutions. That means the model becomes more concise (i.e., faster and cheaper), but it may lose the ability to benefit from additional inference-time compute and can thus perform poorly.
图 29:所提出的 Toggle 方法使 Kimi K2.5 在保持整体基准性能相似的同时,token 效率大大提高。注释图来自 https://arxiv.org/abs/2602.02276
Figure 29: The proposed Toggle method makes Kimi K2.5 much more token-efficient while keeping the overall benchmark performance similar. Annotated figure from https://arxiv.org/abs/2602.02276
Kimi K2.5 的方法称为 Toggle,它在每固定次数的训练迭代中交替进行两个强化学习阶段:
Kimi K2.5’s method, called Toggle, alternates between two RL phases every fixed number of training iterations:
1. 在预算阶段,鼓励正确的解保持在特定于问题的 token 预算之内。
1. In the budgeted phase, correct solutions are encouraged to stay within a problem-specific token budget.
2. 在无约束阶段,恢复通常的最大生成长度,以便模型仍能从较长的解决方案中学习。
2. In the unconstrained phase, the usual maximum generation length is restored so that the model can still learn from longer solutions.
对于每个问题,预算从 RLVR 中正确 rollout 的响应长度的选定百分位数估计。然后,只有当该问题的平均准确率超过阈值时,预算约束才会被激活。这避免了在模型能够可靠地解决问题之前强迫其缩短推理过程。
For each problem, the budget is estimated from a selected percentile of response lengths among correct rollouts in RLVR. The budget constraint is then only activated once the mean accuracy on that problem exceeds a threshold. This avoids forcing the model to shorten its reasoning before it can solve the problem reliably.
图 30:Toggle 方法两个阶段的概览。
Figure 30: Overview of the two phases of the Toggle method.
报告在 K2 Thinking 上评估了 Toggle,发现它减少了约 25% 到 30% 的生成 token,而基准性能几乎没有变化。同样的行为也从数学和编码 RL 任务迁移到了 GPQA 和 MMLU-Pro。
The report evaluates Toggle on K2 Thinking and finds that it reduces generated tokens by about 25 to 30% with little change in benchmark performance. The same behavior also transfers from math and coding RL tasks to GPQA and MMLU-Pro.
Toggle 提供了一种具体的前沿模型训练方案,用于训练更节省 token 的推理策略,同时保持其在测试时扩展的能力。
Toggle supplies a concrete flagship-model recipe for training a more token-efficient reasoning policy while preserving its ability to scale at test time.
Toggle 完全在强化学习训练期间运行。交替的两个阶段更新同一个策略(即 LLM),最终的(统一)检查点没有预算受限与不受限的选择器。在推理时,得到的模型默认以思考模式运行。
Toggle operates entirely during RL training. Both alternating phases update the same policy (i.e., LLM), and the final (unified) checkpoint has no budgeted-versus-unconstrained selector. At inference, the resulting model then runs in thinking mode by default.
有趣的是,在我检查的一些 API(如 vLLM 或 SGLang)中,Kimi K2.5 本身暴露了思考模式和即时模式之间的二元选择。思考模式默认启用。即时模式通过思考禁用推理轨迹:在官方 API 中为 {"type": "disabled"},或通过 vLLM 或 SGLang 服务模型时使用 chat_template_kwargs={"thinking": False}。然而,这些设置与 Toggle 是分开的。
Interestingly, though, Kimi K2.5 itself exposes a separate binary choice between thinking and instant modes in some APIs I checked (like vLLM or SGLang). Thinking mode is enabled by default. Instant mode disables the reasoning trace through thinking: {"type": "disabled"} in the official API or chat_template_kwargs={"thinking": False} when serving the model through vLLM or SGLang. However, these settings are separate from Toggle.
此外,官方 Kimi 报告没有提供即时模式的单独训练配方。然而,K2.5 的 SFT 数据是使用早期的 K2 模型(产生直接响应,没有长推理)和 K2 Thinking(产生扩展推理轨迹)生成的。这可能使统一检查点暴露于两种响应格式,类似于上面 Nemotron 3 的做法。在推理时,聊天模板通过预填充开放的 <think> 标签(思考模式)或空的 <think></think> 块(即时模式)来选择。但同样,不幸的是,报告没有披露确切的数据混合或是否使用了额外的模式特定 RL。
Also, the official Kimi report does not provide a separate training recipe for instant mode. However, K2.5's SFT data were generated using both the earlier K2 model, which produces direct responses without long reasoning, and K2 Thinking, which produces extended reasoning traces. This likely exposes the unified checkpoint to both response formats similar to what's done in Nemotron 3 above. At inference time, the chat template selects between them by prefilling either an open <think> tag for thinking mode or an empty <think></think> block for instant mode. But again, unfortunately, the report does not disclose the exact data mixture or whether additional mode-specific RL was used.
更新的 Kimi K3 提供了更直接的推理时算力接口。当前的 Kimi Code 文档列出了三个设置:low、high 和 max,默认是 max。这些通过 reasoning_effort 参数传递。然而,Moonshot 尚未解释这三个算力级别在训练期间是如何创建的。其发布文章称这些细节将出现在未来的 K3 技术报告中,所以我会持续关注。
The newer Kimi K3 provides a more direct inference-time effort interface. The current Kimi Code documentation lists three settings called low, high, and max, with max as the default. These are passed through the reasoning_effort parameter. However, Moonshot has not yet explained how the three effort levels were created during training. Its launch post says that these details will appear in a future K3 technical report, so I'll stay tuned for that.
根据新的 Kimi K3 技术报告(在我发表本文之后发布),Kimi K3 具有三种推理模式:低强度、高强度和最大强度。
Based on the new Kimi K3 technical report (released after I published this article), Kimi K3 has three reasoning modes: low-, high-, and max-effort.
对于每个训练问题,研究人员首先估计一个初始的 token 预算。在一个简单的可验证任务上,正确的回答通常获得 +1 的奖励,错误的回答获得 0。但在这里,如果回答超过了选定的 token 预算,其奖励则被设为 -1。这给了模型一个强烈的激励,使其保持在预算之内。(对于不可验证的任务和智能体式任务,确切的奖励设置有所不同,Kimi 还使用了学习到的奖励模型和成对比较。)
For each training problem, the researchers first estimate an initial token budget. On a simple verifiable task, a correct response would normally receive a reward of +1 and an incorrect response 0. But here, if the response exceeds the selected token budget, its reward is instead set to -1. This gives the model a strong incentive to stay within the budget. (The exact reward setup is different for non-verifiable and agentic tasks, and Kimi also uses learned reward models and pairwise comparisons.)
为了让 Kimi K3 支持三种推理强度模式,训练过程如下。首先,研究人员从训练一个最大强度版本(“专家”)开始,使用相对宽松的预算。然后他们降低预算来训练高强度和低强度版本。
To have Kimi K3 support the three reasoning effort modes, the training looked as follows. First, the researchers start by training a max-effort version (“specialist”) with a relatively generous budget. They then lower the budget to train high- and low-effort versions.
这个过程在三个领域重复进行:通用任务、通用智能体和编码智能体。每个领域在每个强度级别上都有一个专家。因此,例如,一个专家学习低强度的编码行为,而另一个学习最大强度的智能体行为。这样总共得到九个专家模型。
This process is repeated for three domains: general tasks, general agents, and coding agents. And each domain gets one specialist for each effort level. So, for example, one specialist learns low-effort coding behavior, while another learns max-effort agent behavior. This gives nine specialist models in total.
然后,这九个专家模型的行为通过多教师在线蒸馏合并到一个单一的 Kimi K3 模型中。在推理时,提示中的自然语言思考强度指令可以用来告诉模型使用哪个强度级别。
The behaviors of these nine specialist models are then combined into a single Kimi K3 model using multi-teacher on-policy distillation. At inference time, a natural-language thinking-effort instruction in the prompt can then be used to tell the model which effort level to use.
GLM-5 技术报告将 GLM-4.5 中引入的二元开/关思考开关扩展到多轮和工具使用场景。它描述了三种相关行为(而非三种努力级别):
The GLM-5 technical report extends the binary on/off thinking switch introduced with GLM-4.5 to multi-turn and tool-using scenarios. It describes three related behaviors (rather than three effort levels):
* **交错思考**:在每次响应和工具调用之前插入一个推理块。 * **保留思考**:对话在轮次之间保留早期的推理块,以便模型后续重用。 * **回合级思考**:在对话中为每个请求单独启用或禁用推理。
* Interleaved thinking: this inserts a reasoning block before each response and tool call. * Preserved thinking: here, the chat retains earlier reasoning blocks across turns so that the model can reuse them later. * Turn-level thinking: this enables or disables reasoning separately for each request in a conversation.
在推理时,回合级思考是实际的开关。在 Z.ai API 中,思考默认启用,可通过 thinking: {"type": "disabled"} 为单个请求禁用。托管实现未公开,但开源的 GLM-5 聊天模板展示了在使用 Transformers、vLLM 或 SGLang 自托管时的等效机制。
At inference time, turn-level thinking is the actual on-off switch. In the Z.ai API, thinking is enabled by default and can be disabled for an individual request with thinking: {"type": "disabled"}. The hosted implementation is not disclosed but the open GLM-5 chat template shows the equivalent mechanism when self-hosting with Transformers, vLLM, or SGLang.
The[](https://www.google.com/url?q=https://arxiv.org/abs/2602.15763&sa=D&source=editors&ust=1784335866763244&usg=AOvVaw0AfAPDCBRwHIqgVtzEIlWl)GLM-5 technical report extends the binary on/off thinking switch introduced with GLM-4.5 to multi-turn and tool-using scenarios. It describes three related behaviors (rather than three effort levels):
* Interleaved thinking:this inserts a reasoning block before each response and tool call.
当启用思考时,它以 <|assistant|><think> 开始助手的响应;当禁用思考时,则以 <|assistant|></think> 开始。后者立即关闭推理块,因此生成直接进入最终答案。
It starts the assistant response with <|assistant|><think> when thinking is enabled and <|assistant|></think> when it is disabled. The latter closes the reasoning block immediately, so generation proceeds directly to the final answer.
报告称,这些行为是在多任务 SFT 期间与更新的聊天模板一起引入的。
The report says that these behaviors are introduced during multi-task SFT together with an updated chat template.
在 SFT 之后,GLM-5 依次经历推理强化学习、智能体式强化学习和通用强化学习。最后一步采用在线策略蒸馏,使用前几个阶段的检查点作为教师模型。这有助于最终模型恢复在顺序强化学习阶段可能减弱的能力。
After SFT, GLM-5 goes through reasoning RL, agentic RL, and general RL. And a final on-policy distillation step uses checkpoints from the preceding stages as teachers. This helps the final model recover capabilities that may have weakened during the sequential RL stages.
Qwen3 已在第 4 节中介绍过,因此这里仅总结与本次比较相关的部分。根据 Qwen3 技术报告,其后训练流程包含四个阶段:长思维链 SFT、推理强化学习、思维模式融合和通用强化学习。
Qwen3 was already covered in Section 4, so I will only summarize the parts that matter for this comparison. According to the Qwen3 technical report, its post-training pipeline has four stages. These are long-chain-of-thought SFT, reasoning RL, Thinking Mode Fusion, and general RL.
思维模式融合是实现“努力开关”的关键阶段。在此阶段,模型通过 SFT 在思考与非思考示例的混合数据上进行训练。/think 示例包含推理轨迹,而 /no_think 示例以空的 <think></think> 块开头,并附带简短答案。后续的通用强化学习阶段强化了两种行为的指令遵循和格式遵循。
Thinking Mode Fusion is the key stage for the effort on-off switch. Here the model is trained via SFT on a mixture of thinking and non-thinking examples. The /think examples contain a reasoning trace, while /no_think examples begin with an empty <think></think> block that is accompanied by a short answer. The following general RL stage reinforces instruction and format following for both behaviors.
Qwen3 还支持硬性思考预算。在达到请求的阈值时,推理跨度被停止,并在模型继续生成最终答案之前插入停止思考指令。报告指出,这种部分推理行为并非显式训练出来的,而是在思维模式融合之后涌现的。
Qwen3 also supports a hard thinking budget. At the requested threshold, the reasoning span is stopped and a stop-thinking instruction is inserted before the model continues with its final answer. The report says that this partial-reasoning behavior was not trained explicitly. It emerged after Thinking Mode Fusion.
这为 Qwen3 提供了学习到的开关以及推理时预算。它与 DeepSeek V4 和 Nemotron 的方案类似,但更为简单。
This gives Qwen3 a learned on-off switch plus an inference-time budget. It is similar but simpler than the DeepSeek V4 and Nemotron recipes.
Inkling 已在第 5.3 节讨论过。简而言之,其技术报告提到他们使用连续努力条件(值在 0.0 到 1.0 之间),而不是固定的努力标签。
Inkling was already discussed in Section 5.3. The short version is that its technical report mentions that they use continuous effort conditioning (values between 0.0 and 1.0) rather than fixed effort labels.
在相对较小的初始 SFT 阶段之后,Inkling 的大部分后训练来自异步强化学习,超过 3000 万次 rollout。期望的努力值包含在系统消息中,并在强化学习期间根据该值调整 token 长度惩罚。如前所述,较高的 token 成本鼓励较短的响应,较低的 token 成本给模型更多推理空间。
After a relatively small initial SFT stage, most of Inkling's post-training comes from asynchronous RL with more than 30 million rollouts. The desired effort is included in the system message, and the token length penalty is adjusted according to that value during RL. As previously discussed, a higher token cost encourages a shorter response. A lower token cost gives the model more room to reason.
下表总结了六份技术报告中实际记录的内容。
The table below summarizes what is actually documented in the six technical reports.
图 32:六个具有推理努力设置的开放权重模型所披露的训练机制和推理控制的比较。
Figure 32: Comparison of the disclosed training mechanisms and inference controls for six open-weight models with reasoning-effort settings.
因此,观察这六个不同的开放权重模型,它们有一个共同的框架。首先,它们通过 SFT 和聊天模板引入努力模式控制。Qwen3 明确混合了思考和非思考示例,而 GLM-5 则增加了交错、保留和回合级思考模式。
So, looking at the six different open-weight models, they have a shared framework. First, they introduce effort mode control through SFT and the chat template. Qwen3 explicitly mixes thinking and non-thinking examples, while GLM-5 adds interleaved, preserved, and turn-level thinking patterns.
第二个共同组件是模式条件强化学习阶段,其中上下文窗口和长度惩罚随请求的努力程度而变化。DeepSeek V4、Nemotron 3 Ultra 和 Inkling 采用了这种方法。
The second shared component is a mode-conditioned RL stage, where context windows and length penalties change with the requested effort. DeepSeek V4, Nemotron 3 Ultra, and Inkling use this approach.
第三个要素是在显式预算下提高鲁棒性。Nemotron 在随机截断的轨迹上训练,Qwen3 可以从强制停止的推理片段继续,而 Kimi 则将预算约束与无约束的强化学习交替进行。这些方法有助于在可用推理长度变化甚至被截断时保持答案质量。
A third ingredient improves robustness under explicit budgets. Nemotron trains on randomly truncated traces, Qwen3 can continue from a forcibly stopped reasoning span, and Kimi alternates budgeted with unconstrained RL. These methods help preserve answer quality when the available reasoning length changes and is even cut short.
本文中的开放权重示例通过几种不同的机制来实现推理努力(reasoning effort)。相似的标签可能由独立的专家模型、混合的 SFT 数据、模式条件奖励、硬性 token 预算或这些方法的组合来支撑。
The open-weight examples in this article implement reasoning effort through several different mechanisms. Similar labels can be backed by separate specialists, mixed SFT data, mode-conditioned rewards, hard token budgets, or combinations of these methods.
很难说哪种方法最好。这些模型在基础检查点、训练数据、后训练算力、基准测试和服务目标上各不相同。它们的报告也省略了许多受控比较所需的细节。(此外,可能不存在一种万能的方法,对交互式助手有效的方法可能不适合长时间运行的编码智能体。)
It is difficult to say which approach is best. The models differ in their base checkpoints, training data, post-training compute, benchmarks, and serving goals. Their reports also omit many details needed for a controlled comparison. (Also, there may not be a one-size-fits-all, and a method that works well for an interactive assistant may be a poor fit for a long-running coding agent.)
圣杯当然是自动努力选择。我们之前在 GPT 5 的 Auto 模式中见过这一点。这是一个棘手的问题,最终实现可能弊大于利,这就是它从界面中移除的原因(至少我再也找不到它了)。
The holy grail is of course automatic effort selection. We saw this a while back with GPT 5's Auto mode. It's a tricky problem to solve, and in the end, the implementation was probably more miss than hit, which is why it got removed from the UI (at least, I can't find it anymore).
在不久的将来,我认为推理努力仍将是一个显式的模型输入,通常通过系统提示来传递。然而,围绕 LLM 的智能体包装器/框架,或内部路由器,可能会越来越多地从任务状态和可用资源中自动推断适当的模式和预算(当然仍然允许用户覆盖)。
In the near future, I think reasoning effort will remain an explicit model input, which will most often be delivered through the system prompt. However agent wrapper/harness around the LLM, or an internal router may increasingly infer the appropriate mode and budget from the task state and available resources automatically (while of course still allowing a user override).
我仍然希望努力选择能变得更加自动化。类似于 GPT 5 的自动模式,一个廉价的模型或路由器可以根据请求、工具状态以及剩余时间或 token 预算来选择模式,同时仍然允许用户覆盖。如果你想优化延迟、成本或最大性能,覆盖功能是很有用的。
I still hope that effort selection will become more automatic. Similar to GPT 5's auto mode, a cheap model or router could choose the mode from the request, tool state, and remaining time or token budget while still allowing a user override. The override is useful if you want to optimize for latency or cost, or maximum performance.
我意识到这是一篇很长的文章,而且可能不是最引人注目的主题。但考虑到目前关于 LLM、推理模型和智能体的讨论很多,我认为对推理模型进行探讨是之前未曾涉及的,希望这是一份独特且有一定用处的概述!
I realize that this was a long article, and it was perhaps not the flashiest topic. But I thought that given all the talk about LLMs, reasoning models, and agents, a look at reasoning models was something not covered before, and I hope it was a unique and somewhat useful overview!
如果你想动手实现推理模型背后的核心训练方法,我的《从零构建推理模型》一书会逐步讲解带可验证奖励的强化学习和推理时扩展,并附有代码。
If you want a hands-on implementation of the core training methods behind reasoning models, my *Build a Reasoning Model (From Scratch)* book walks through reinforcement learning with verifiable rewards and inference-time scaling step by step, with code.
本文聚焦于训练好的推理模型如何支持不同的努力模式。而这本书则退一步,首先展示如何将传统大语言模型转变为推理模型。它是《从零构建大语言模型》的续作,从上一本书结束的地方开始。
This article focused on how a trained reasoning model can support different effort modes. The book takes a step back and shows how to turn a conventional LLM into a reasoning model in the first place. It is a sequel to *Build a Large Language Model (From Scratch)* and starts where that book leaves off.
印刷版现已开始发货。
The print edition has now started shipping.
《从零构建推理模型》[Manning] [Amazon]
*Build a Reasoning Model (From Scratch)* [Manning] [Amazon]
如果你喜欢我之前写的《构建大语言模型(从零开始)》一书,那么这本质上是一本续作,从零开始实现推理时缩放技术和强化学习算法。
If you liked my previous Build a Large Language Model (From Scratch) book, this is essentially a sequel implementing inference-time scaling techniques and reinforcement learning algorithms from scratch.
如果你想支持未来像这样的长篇文章,可以考虑成为付费订阅者。这有助于我继续撰写这些独立的深度文章,并分享随附的代码、图表和实验。
And if you want to support future long-form articles like this one, consider becoming a paid subscriber. It helps me keep writing these independent deep dives and sharing the accompanying code, figures, and experiments.