Why We Think
打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→日期:2025 年 5 月 1 日 | 预计阅读时间:40 分钟 | 作者:Lilian Weng * 模型是否忠实反映其思考 * 思维链上的优化压力:好还是坏?
Date: May 1, 2025 | Estimated Reading Time: 40 min | Author: Lilian Weng * Does the Model Tell What it Thinks Faithfully * Optimization Pressure on CoT: Good or Bad?
日期:2025 年 5 月 1 日 | 预计阅读时间:40 分钟 | 作者:Lilian Weng
Date: May 1, 2025 | Estimated Reading Time: 40 min | Author: Lilian Weng
* 模型是否忠实地表达了它的想法
* Does the Model Tell What it Thinks Faithfully
* 对思维链的优化压力:好还是坏?
* Optimization Pressure on CoT: Good or Bad?
特别感谢 John Schulman 对本文的大量宝贵反馈和直接编辑。
Special thanks to John Schulman for a lot of super valuable feedback and direct edits on this post.
测试时算力(Graves 等人,2016;Ling 等人,2017;Cobbe 等人,2021)和思维链(Wei 等人,2022;Nye 等人,2021)显著提升了模型性能,同时也引发了许多研究问题。本文旨在回顾如何有效利用测试时算力(即“思考时间”)及其有效原因的最新进展。
Test time compute (Graves et al. 2016, Ling, et al. 2017, Cobbe et al. 2021) and Chain-of-thought (CoT) (Wei et al. 2022, Nye et al. 2021), have led to significant improvements in model performance, while raising many research questions. This post aims to review recent developments in how to effectively use test-time compute (i.e. “thinking time”) and why it helps.
让模型进行更长时间的思考可以从几个不同的角度来论证。
Enabling models to think for longer can be motivated in a few different ways.
核心思想与人类思考方式密切相关。我们人类无法立即回答“12345 乘以 56789 等于多少?”相反,对于复杂问题,我们自然会花时间思考和分析才能得出结果。在《思考,快与慢》(Kahneman, 2013)中,丹尼尔·卡尼曼通过双过程理论将人类思维分为两种模式:
The core idea is deeply connected to how humans think. We humans cannot immediately provide the answer for "What's 12345 times 56789?". Rather, it is natural to spend time pondering and analyzing before getting to the result, especially for complex problems. In Thinking, Fast and Slow (Kahneman, 2013), Daniel Kahneman characterizes human thinking into two modes, through the lens of the dual process theory :
* _快思考(系统 1)_ 运行迅速且自动,由直觉和情感驱动,几乎不需要努力。
* _Fast thinking (System 1)_ operates quickly and automatically, driven by intuition and emotion while requiring little to no effort.
* _慢思考(系统 2)_ 需要深思熟虑、逻辑思维和大量的认知努力。这种思维模式消耗更多脑力,需要有意投入。
* _Slow thinking (System 2)_ demands deliberate, logical thought and significant cognitive efforts. This mode of thinking consumes more mental energy and requires intentional engagement.
由于系统 1 思考快速且轻松,它往往成为主要的决策驱动者,但代价是准确性和逻辑性。它自然依赖于大脑的心理捷径(即启发式),可能导致错误和偏见。通过有意识地放慢速度,花更多时间反思、改进和分析,我们可以启动系统 2 思考,挑战直觉并做出更理性的选择。
Because System 1 thinking is fast and easy, it often ends up being the main decision driver, at the cost of accuracy and logic. It naturally relies on our brain’s mental shortcuts (i.e., heuristics) and can lead to errors and biases. By consciously slowing down and taking more time to reflect, improve and analyze, we can engage in System 2 thinking to challenge our instincts and make more rational choices.
深度学习的一种观点是,神经网络可以通过前向传播中能够访问的算力和存储量来表征,如果我们通过梯度下降优化它们来解决问题,优化过程将找出如何使用这些资源——它们将找出如何将这些资源组织成用于计算和信息存储的电路。从这个角度看,如果我们设计一个在测试时能够进行更多计算的架构或系统,并训练它有效利用这一资源,那么它的表现会更好。
One view of deep learning, is that neural networks can be characterized by the amount of computation and storage they can access in a forward pass, and if we optimize them to solve problems using gradient descent, the optimization process will figure out how to use these resources–they’ll figure out how to organize these resources into circuits for calculation and information storage. From this view, if we design an architecture or system that can do more computation at test time, and we train it to effectively use this resource, it’ll work better.
在 Transformer 模型中,模型为每个生成的词元所执行的计算量(flops)大约是参数数量的两倍。对于混合专家(MoE)这样的稀疏模型,每次前向传播中只使用一部分参数,因此计算量 = 2 * 参数数量 / 稀疏度,其中稀疏度是活跃专家的比例。
In Transformer models, the amount of computation (flops) that the model does for each generated token is roughly 2 times the number of parameters. For sparse models like mixture of experts (MoE), only a fraction of the parameters are used in each forward pass, so computation = 2 * parameters / sparsity, where sparsity is the fraction of experts active.
另一方面,思维链使模型能够为它试图计算的每个答案词元执行多得多的 flops 计算。事实上,思维链有一个很好的特性,即它允许模型根据问题的难度使用可变数量的计算资源。
On the other hand, CoT enables the model to perform far more flops of computation for each token of the answer that it is trying to compute. In fact, CoT has a nice property that it allows the model to use a variable amount of compute depending on the hardness of the problem.
机器学习中的一个经典思想是定义一个包含潜在(隐藏)变量 $z$ 和可见变量 $y$ 的概率模型,其中 $y$ 提供给我们的学习算法。通过对潜在变量的可能取值进行边缘化(求和),我们可以表达可见变量上的丰富分布:$P(y) = \sum_{z \sim P(z)} P(y \mid z)$。例如,我们可以通过令 $x$ 表示问题陈述,$y$ 表示真实答案或证明,$z$ 表示导致证明的自由形式思维过程,来建模数学问题和解的分布。要优化的边缘概率分布为 $P(y \mid x) = \sum_{z \sim p(z \mid x)} P(y \mid x, z)$。
A classic idea in machine learning is to define a probabilistic model with a latent (hidden) variable $z$ and a visible variable $y$, where $y$ is given to our learning algorithm. Marginalizing (summing) over the possible values of the latent variable allows us to express a rich distribution over the visible variables, $P \left(\right. y \left.\right) = \underset{z sim P \left(\right. z \left.\right)}{\sum} P \left(\right. y \mid z \left.\right)$. For example, we can model the distribution over math problems and solutions by letting $x$ denote a problem statement, $y$ be ground truth answer or proof, and $z$ as a free-form thought process that leads to the proof. The marginal probability distribution to optimize would be $P \left(\right. y \mid x \left.\right) = \underset{z sim p \left(\right. z \mid x \left.\right)}{\sum} P \left(\right. y \mid x , z \left.\right)$
潜在变量视角对于理解涉及收集多个并行思维链或搜索思维链的方法特别有用——这些算法可以看作是从后验分布 $P(z \mid x, y)$ 中采样。这一观点也表明使用对数损失 $\log P(y \mid x)$ 作为优化目标的好处,因为对数损失目标在预训练中非常有效。
The latent variable perspective is particularly useful for understanding methods that involve collecting multiple parallel CoTs or searching over the CoT–these algorithms can be seen as sampling from the posterior $P \left(\right. z \mid x , y \left.\right)$. This view also suggests the benefits of using the log loss $log P \left(\right. y \mid x \left.\right)$ as the target objective to optimize, as the log loss objective has been so effective in pretraining.
在生成简短答案之前先生成中间步骤的策略,特别是针对数学问题,由 Ling 等人(2017)探索,他们引入了 AQUA-RAT 数据集,随后由 Cobbe 等人(2021)扩展,引入了小学数学(GSM)数据集。Cobbe 等人通过监督学习在人工编写的解决方案上训练生成器,并训练验证器预测候选解决方案的正确性;然后他们可以搜索这些解决方案。Nye 等人(2021)实验了将中间思考 token 作为“草稿板”,而 Wei 等人(2022)提出了现在标准的术语“思维链”(CoT)。
The strategy of generating intermediate steps before generating short answers, particularly for math problems, was explored by Ling, et al. 2017, who introduced the AQUA-RAT dataset, and then expanded by Cobbe et al. 2021, who introduced the Grade School Math (GSM) dataset. Cobbe et al. train a generator with supervised learning on human-written solutions and verifiers that predict the correctness of a candidate solution; they can then search over these solutions. Nye et al. (2021) experimented with intermediate thinking tokens as “scratchpads” and Wei et al. (2022) coined the now-standard term chain-of-thought (CoT).
早期改进 CoT 推理的工作包括在人工编写的推理轨迹或经过答案正确性筛选的模型编写轨迹上进行监督学习,后者可视为强化学习(RL)的初级形式。其他一些工作发现,通过适当的提示,如“逐步思考”(Kojima 等人,2022)或更复杂的提示以鼓励模型先反思相关知识(Yasunaga 等人,2023),可以显著提升指令调优模型的数学性能。
Early work on improving CoT reasoning involved doing supervised learning on human-written reasoning traces or model-written traces filtered for answer correctness, where the latter can be seen as a rudimentary form of reinforcement learning (RL). Some other work found that one could significantly boost math performance of instruction tuned models by prompting them appropriately, with "think step by step" (Kojima et al. 2022) or more complex prompting to encourage the model to reflect on related knowledge first (Yasunaga et al. 2023).
后来的工作发现,通过在具有自动可检查解决方案的问题数据集上进行强化学习,例如具有简短答案的 STEM 问题或可通过单元测试检查的编码任务,可以显著提升 CoT 推理能力(Zelikman 等人,2022;Wang 等人,2023;Liu 等人,2023)。这种方法随着 o1-preview、o3 的发布以及 R1 技术报告(DeepSeek-AI,2025)而声名鹊起,该报告展示了一个简单的策略梯度算法配方即可带来强大的性能。
Later work found that the CoT reasoning capabilities can be significantly improved by doing reinforcement learning on a dataset of problems with automatically checkable solutions, such as STEM problems with short answers, or coding tasks that can be checked with unit tests (Zelikman et al. 2022, Wang et al., 2023, Liu et al., 2023). This approach rose to prominence with the announcement of o1-preview, o3, and the R1 tech report (DeepSeek-AI, 2025), which showed that a simple recipe where a policy gradient algorithm could lead to strong performance.
思维链提示提高了解决数学问题的成功率。更大的模型从思考时间中获益更多。(图片来源:Wei 等人,2022)
Chain-of-thought prompting leads to higher success rate of solving math problems. Larger models benefit more from thinking time. (Image source: Wei et al. 2022)
测试时算力的基本意图是在测试时自适应地修改模型的输出分布。有多种利用测试时资源进行解码的方法,以选择更好的样本,从而将模型的预测向更期望的分布改变。改进解码过程的两种主要方法是并行采样和顺序修订。
The fundamental intent of test-time compute is to adaptively modify the model’s output distribution at test time. There are various ways of utilizing test time resources for decoding to select better samples and thus alter the model’s predictions towards a more desired distribution. Two main approaches for improving the decoding process are parallel sampling and sequential revision.
* 并行采样同时生成多个输出,同时在每一步使用过程奖励信号提供指导,或在最后使用验证器判断质量。这是最广泛采用的提高测试时性能的解码方法,例如 best-of-$N$ 或束搜索。自一致性(Wang 等人,2023)常用于在无法获得真实答案时,通过多个思维链展开的多数投票选择答案。
* Parallel sampling generates multiple outputs simultaneously, meanwhile providing guidance per step with process reward signals or using verifiers to judge the quality at the end. It is the most widely adopted decoding method to improve test time performance, such as best-of-$N$ or beam search. Self-consistency (Wang et al. 2023) is commonly used to select the answer with majority vote among multiple CoT rollouts when the ground truth is not available.
* 顺序修订基于前一步的输出迭代调整模型的响应,要求模型有意识地反思现有响应并纠正错误。修订过程可能依赖于微调后的模型,因为单纯依赖模型内在的自我纠正能力(无外部反馈)可能不会带来改进(Kamoi 等人,2024;Huang 等人,2024)。
* Sequential revision adapts the model’s responses iteratively based on the output in the previous step, asking the model to intentionally reflect its existing response and correct mistakes. The revision process may have to rely on a fine-tuned model, as naively relying on the model’s intrinsic capability of self-correction without external feedback may not lead to improvement (Kamoi et al. 2024, Huang et al. 2024).
并行采样简单、直观且易于实现,但受限于模型能否一次性达到正确解的能力。顺序修订明确要求模型反思错误,但速度较慢,且在实现时需要额外小心,因为它存在将正确预测修改为错误或引入其他类型幻觉的风险。这两种方法可以结合使用。Snell 等人(2024)表明,简单问题受益于纯顺序的测试时算力,而困难问题通常在顺序与并行算力的最优比例下表现最佳。
Parallel sampling is simple, intuitive and easier to implement, but bounded by the model capability of whether it can achieve the correct solution in one-go. Sequential explicitly asks the model to reflect on mistakes but it is slower and requires extra care during implementation as it does run the risk of correct predictions being modified to be incorrect or introducing other types of hallucinations. These two methods can be used together. Snell et al. (2024) showed that easier questions benefit from purely sequential test-time compute, whereas harder questions often perform best with an optimal ratio of sequential to parallel compute.
并行采样与顺序修订的示意图。
Illustration of parallel sampling vs sequential revision.
给定一个生成模型和一个可用于对完整或部分样本进行评分的评分函数,我们可以使用多种搜索算法来寻找高分样本。Best-of-$N$ 是最简单的此类算法:只需收集 $N$ 个独立样本,并根据某个评分函数选择排名最高的样本。束搜索是一种更复杂的搜索算法,它使搜索过程更具适应性,将更多采样算力投入到解空间中更有希望的部分。
Given a generative model and a scoring function that we can use to score full or partial samples, there are various search algorithms we can use to find a high scoring sample. Best-of-$N$ is the simplest such algorithm: one just collects $N$ independent samples and chooses the highest-ranking sample according to some scoring function. Beam search is a more sophisticated search algorithm that makes the search process more adaptive, spending more sampling computation on more promising parts of the solution space.
束搜索维护一组有希望的部分序列,并在扩展它们和修剪较不有希望的序列之间交替。作为选择机制,我们可以使用过程奖励模型(PRM;Lightman 等人,2023)来指导束搜索候选选择。Xie 等人(2023)使用 LLM 评估其自身生成的推理步骤正确的可能性,格式化为多项选择题,并发现每步自我评估减少了束搜索解码过程中多步推理的累积误差。此外,在采样过程中,退火温度有助于减轻聚合随机性。Xie 等人的这些实验在使用 Codex 模型的少样本 GSM8k、AQuA 和 StrategyQA 基准测试上实现了 5-6% 的提升。奖励平衡搜索(简称“REBASE”;Wu 等人,2025)单独训练了一个过程奖励模型(PRM),根据 softmax 归一化的奖励分数,确定在束搜索过程中每个深度应扩展每个节点多少。Jiang 等人(2024)训练了他们的 PRM,名为“RATIONALYST”,用于在大量未标注数据上基于合成理由进行束搜索指导。好的理由根据它们是否有助于将真实答案 token 的负对数概率降低一个阈值来筛选,通过比较理由包含在上下文中与不包含时的差异。在推理时,RATIONALYST 通过帮助估计下一步推理步骤的对数概率(“隐式”)或直接生成下一步推理步骤作为提示的一部分(“显式”),为 CoT 生成器提供过程监督。
Beam search maintains a set of promising partial sequences and alternates between extending them and pruning the less promising ones. As a selection mechanism, we can use a process reward model (PRM; Lightman et al. 2023) to guide beam search candidate selection. Xie et al. (2023) used LLM to evaluate how likely its own generated reasoning step is correct, formatted as a multiple-choice question and found that per-step self-evaluation reduces accumulative errors in multi-step reasoning during beam search decoding. Besides, during sampling, annealing the temperature helps mitigate aggregated randomness. These experiments by Xie et al. achieved 5-6% improvement on few-shot GSM8k, AQuA and StrategyQA benchmarks with the Codex model. Reward balanced search (short for “REBASE”; Wu et al. 2025) separately trained a process reward model (PRM) to determine how much each node should be expanded at each depth during beam search, according to the softmax-normalized reward scores. Jiang et al. (2024) trained their PRM, named “RATIONALYST”, for beam search guidance on synthetic rationales conditioned on a large amount of unlabelled data. Good rationales are filtered based on whether they help reduce the neg log-prob of true answer tokens by a threshold, when comparing the difference between when the rationales is included in the context vs not. At inference time, RATIONALYST provides process supervision to the CoT generator by helping estimate log-prob of next reasoning steps (“implicit”) or directly generating next reasoning steps as part of the prompt (“explicit”).
由 LLM 每步自我评估引导的束搜索解码。(图片来源:Xie 等人,2023)
Beam search decoding guided by LLM self-evaluation per reasoning step. (Image source: Xie et al. 2023)
有趣的是,可以在没有显式零样本或少样本提示的情况下触发涌现的思维链推理路径。Wang & Zhou(2024)发现,如果我们在第一个采样 token 处进行分支,保留置信度最高的前 $k$ 个 token(置信度通过采样时 top-1 和 top-2 候选之间的差异衡量),然后对这些 $k$ 个采样试验继续使用贪心解码,许多这样的序列自然包含 CoT。特别是当 CoT 出现在上下文中时,它会导致对最终答案的更自信解码。为了计算最终答案的置信度,需要通过任务特定的启发式方法(例如,数学问题的最后一个数值)或通过进一步提示模型“所以答案是”来识别答案跨度。仅在第一个 token 处进行分支的设计选择基于这样的观察:早期分支显著增加了潜在路径的多样性,而后续 token 受先前序列影响较大。
Interestingly, it is possible to trigger the emergent chain-of-thought reasoning paths _without_ explicit zero-shot or few-shot prompting. Wang & Zhou (2024) discovered that if we branch out at the first sampling tokens by retaining the top $k$ tokens with highest confidence, measured as the difference between top-1 and top-2 candidates during sampling, and then continue these $k$ sampling trials with greedy decoding onward, many of these sequences natively contain CoT. Especially when CoT does appear in the context, it leads to a more confident decoding of the final answer. To calculate the confidence of the final answer, the answer span needs to be identified by task-specific heuristics (e.g. last numerical values for math questions) or by prompting the model further with "So the answer is". The design choice of only branching out at the first token is based on the observation that early branching significantly enhances the diversity of potential paths, while later tokens are influenced a lot by previous sequences.
Top-$k$ 解码,$k$ 指第一个采样步骤的候选数量。(图片来源:Wang & Zhou,2024)
Top-$k$ decoding, $k$ refers to the number of candidates at the first sampling step. (Image source: Wang & Zhou, 2024)
如果模型能够反思并纠正过去回答中的错误,我们期望模型能产生一系列质量逐步提升的迭代修订。然而,这种自我纠正能力实际上并非大型语言模型(LLM)固有的,并且由于多种失败模式,它并不容易开箱即用,例如:(1) 幻觉,包括将正确的回答修改为错误的;(2) 行为崩溃为不纠正行为,例如对最初的错误回答进行微小修改或不修改;或 (3) 在测试时无法泛化到分布偏移。Huang 等人 (2024) 的实验表明,简单应用自我纠正会导致性能下降,模型需要外部反馈才能自我改进,这些反馈可以基于匹配真实值、启发式方法和任务特定指标、编程问题的单元测试结果 (Shinn 等人, 2023)、更强的模型 (Zhang 等人, 2024) 以及人类反馈 (Liu 等人, 2023)。
If the model can reflect and correct mistakes in past responses, we would expect the model to produce a nice sequence of iterative revision with increasing quality. However, this self-correction capability turns out to not exist intrinsically among LLMs and does not easily work out of the box, due to various failure modes, such as, (1) hallucination, including modifying correct responses to be incorrect; (2) behavior collapse to non-correcting behavior; e.g. making minor or no modification on the first incorrect responses; or (3) fail to generalize to distribution shift at test time. Experiments by Huang et al. (2024) showed that naively applying self-correction leads to worse performance and external feedback is needed for models to self improve, which can be based on matching ground truths, heuristics and task-specific metrics, unit tests results for coding questions (Shinn, et al. 2023), a stronger model (Zhang et al. 2024), as well as human feedback (Liu et al. 2023).
自我纠正学习 (Welleck 等人, 2023) 旨在训练一个纠正器模型 $P_{\theta} \left(\right. y \mid y_{0} , x \left.\right)$,给定一个固定的生成器模型 $P_{0} \left(\right. y_{0} \mid x \left.\right)$。虽然生成器模型保持通用,但纠正器模型可以是任务特定的,并且仅基于初始模型响应和额外反馈(例如,一个句子、编译器跟踪、单元测试结果;可选)进行生成:
Self-correction learning (Welleck et al. 2023) aims to train a corrector model $P_{\theta} \left(\right. y \mid y_{0} , x \left.\right)$ given a fixed generator model $P_{0} \left(\right. y_{0} \mid x \left.\right)$. While the generator model remains to be generic, the corrector model can task-specific and only does generation conditioned on an initial model response and additional feedback (e.g. a sentence, a compiler trace, unit test results; can be optional):
1. 自我纠正学习首先在数据池中为每个提示生成多个输出;
1. Self-correction learning first generates first generates multiple outputs per prompt in the data pool;
2. 然后通过将同一提示的两个输出配对(如果一个输出的价值高于另一个),创建价值改进对(提示 $x$,假设 $y$,纠正 $y^{'}$)。
2. then create value-improving pairs by pairing two outputs for the same prompt together if one has a higher value than the other, (prompt $x$, hypothesis $y$, correction $y^{'}$).
3. 这些对根据其价值改进 $v \left(\right. y^{'} \left.\right) - v \left(\right. y \left.\right)$ 以及两个输出之间的相似度 $\text{Similarity} \left(\right. y , y^{'} \left.\right)$ 进行选择,以训练纠正器模型。
3. These pairs are selected proportional to is improvement in value, $v \left(\right. y^{'} \left.\right) - v \left(\right. y \left.\right)$, and similarity between two outputs, $\text{Similarity} \left(\right. y , y^{'} \left.\right)$ to train the corrector model.
4. 为了鼓励探索,纠正器将新的生成结果也加入数据池。在推理时,纠正器可以迭代使用,以创建顺序修订的纠正轨迹。
4. To encourage exploration, the corrector provides new generations into the data pool as well. At the inference time, the corrector can be used iteratively to create a correction trajectory of sequential revision.
通过匹配同一问题的模型输出形成价值改进对来训练纠正模型的自我纠正学习示意图。(图片来源:Welleck 等人, 2023)
Illustration of self-correction learning by matching model outputs for the same problem to form value-improving pairs to train a correction model. (Image source: Welleck et al. 2023)
递归检查 (Qu 等人, 2024) 也旨在训练更好的纠正器模型,但使用单个模型同时进行生成和自我纠正。
Recursive inspection (Qu et al. 2024) also aims to train a better corrector model but with a single model to do both generation and self-correction.
SCoRe(通过强化学习进行自我纠正;Kumar 等人, 2024)是一种多轮强化学习方法,通过鼓励模型在第二次尝试中产生比第一次尝试更好的答案来进行自我纠正。它包含两个训练阶段:第一阶段仅最大化第二次尝试的准确性,同时对第一次尝试施加 KL 惩罚,以避免第一次尝试的响应过多偏离基础模型行为;第二阶段优化第一次和第二次尝试产生的答案的准确性。理想情况下,我们希望第一次和第二次尝试的性能都更好,但加入第一阶段可以防止模型对第一次响应进行微小或不修改的行为崩溃,第二阶段进一步改进了结果。
SCoRe (Self-Correction via Reinforcement Learning; Kumar et al. 2024) is a multi-turn RL approach to encourage the model to do self-correction by producing better answers at the second attempt than the one created at the first attempt. It composes two stages of training: stage 1 only maximizes the accuracy of the second attempt while enforcing a KL penalty only on the first attempt to avoid too much shifting of the first-turn responses from the base model behavior; stage 2 optimizes the accuracy of answers produced by both the first and second attempts. Ideally we do want to see performance at both first and second attempts to be better, but adding stage 1 prevents the behavior collapse where the model does minor or none edits on the first response, and stage 2 further improves the results.
通过两阶段强化学习训练来提升自我纠正能力的显式训练设置。(图片来源:Kumar 等人, 2024)
Explicit training setup to improve self-correction capabilities by doing two-staged RL training. (Image source: Kumar et al. 2024)
最近,使用强化学习(RL)来提升语言模型推理能力取得了许多成功,其方法是通过收集一系列带有标准答案的问题(通常是 STEM 问题和谜题,答案易于验证),并在模型给出正确答案时给予奖励。该领域近期的活跃是由 OpenAI 的 o 系列模型的强劲表现,以及随后 DeepSeek 发布的模型和技术报告所推动的。
There’s been a lot of recent success in using RL to improve the reasoning ability of language models, by using a collection of questions with ground truth answers (usually STEM problems and puzzles with easy to verify answers), and rewarding the model for getting the correct answer.Recent activity in this area was spurred by strong performance of the o-series models from OpenAI, and the subsequent releases of models and tech reports from DeepSeek.
DeepSeek-R1(DeepSeek-AI, 2025)是一个开源的大语言模型,旨在擅长需要高级推理技能的任务,如数学、编码和逻辑问题解决。他们进行了两轮 SFT-RL 训练,使 R1 在推理和非推理任务上都表现出色。
DeepSeek-R1 (DeepSeek-AI, 2025) is an open-source LLM designed to excel in tasks that require advanced reasoning skills like math, coding and logical problem solving. They run through 2 rounds of SFT-RL training, enabling R1 to be good at both reasoning and non-reasoning tasks.
1. 冷启动 SFT 是在 DeepSeek-V3-Base 基础模型上,使用数千条冷启动数据进行微调。没有这一步,模型会出现可读性差和语言混杂的问题。
1. Cold-start SFT is to fine-tune the DeepSeek-V3-Base base model on a collection of thousands of cold-start data. Without this step, the model has issues of poor readability and language mixing.
2. 面向推理的 RL 仅在推理提示上训练推理模型,使用两种基于规则的奖励:
2. Reasoning-oriented RL trains a reasoning model on reasoning-only prompts with two types of rule-based rewards:
* 格式奖励:模型应将思维链(CoT)包裹在<thinking>...</thinking>标记中。
* Format rewards: The model should wrap CoTs by <thinking> ... </thinking> tokens.
* 准确率奖励:最终答案是否正确。数学问题的答案需要以特定格式(例如,在方框中)呈现以便可靠验证。对于编码问题,使用编译器评估测试用例是否通过。
* Accuracy rewards: Whether the final answers are correct. The answer for math problems needs to be present in a specific format (e.g. in a box) to be verified reliably. For coding problems, a compiler is used to evaluate whether test cases pass.
3. 拒绝采样+非推理 SFT 利用对第 2 步的 RL 检查点进行拒绝采样创建的新 SFT 数据,结合来自 DeepSeek-V3 在写作、事实问答和自我认知等领域的非推理监督数据,重新训练 DeepSeek-V3-Base。
3. Rejection-sampling + non-reasoning SFT utilizes new SFT data created by rejection sampling on the RL checkpoint of step 2, combined with non-reasoning supervised data from DeepSeek-V3 in domains like writing, factual QA, and self-cognition, to retrain DeepSeek-V3-Base.
* 过滤掉语言混杂、长段落和包含代码块的思维链。
* Filter out CoTs with mixed languages, long paragraphs, and code blocks.
* 使用 DeepSeek-V3(DeepSeek-AI, 2024)流程包含非推理任务。
* Include non-reasoning tasks using DeepSeek-V3 (DeepSeek-AI, 2024) pipeline.
* 对于某些非推理任务,在回答问题前通过提示调用 DeepSeek-V3 生成潜在的思维链。但对于像“hello”这样的简单查询,不需要思维链。
* For certain non-reasoning tasks, call DeepSeek-V3 to generate potential CoTs before answering the question by prompting. But for simpler queries like “hello”, CoT is not needed.
* 然后在总共 80 万样本上对 DeepSeek-V3-Base 进行微调,训练 2 个 epoch。
* Then fine-tune the DeepSeek-V3-Base on the total 800k samples for 2 epochs.
4. 最后的 RL 阶段在第 3 步的检查点上同时使用推理和非推理提示进行训练,提升有用性、无害性和推理能力。
4. The final RL stage trains the step 3 checkpoint on both reasoning and non-reasoning prompts, improving helpfulness, harmlessness and reasoning.
DeepSeek-R1 在几个广泛使用的推理基准上表现与 OpenAI o1-preview 和 o1-mini 相当。DeepSeek-V3 是列表中唯一的非推理模型。(图片来源:DeepSeek-AI, 2025)
DeepSeek-R1 performs comparable to OpenAI o1-preview and o1-mini on several widely used reasoning benchmarks. DeepSeek-V3 is the only non-reasoning model listed. (Image source: DeepSeek-AI, 2025)
有趣的是,DeepSeek 团队表明,仅使用 RL,没有 SFT 阶段,仍然可以学习到高级推理能力,如反思和回溯(“顿悟时刻”)。模型在 RL 训练过程中自然地学会了使用更多的思考 token 来解决推理任务。“顿悟时刻”可以出现,指的是模型反思之前的错误,然后尝试替代方法来纠正它们。后来,出现了各种开源努力来复现 R1 的结果,如 Open-R1、SimpleRL-reason 和 TinyZero,它们都基于 Qwen 模型。这些努力也证实了纯 RL 在数学问题上能带来出色的性能,以及“顿悟时刻”的出现。
Interestingly the DeepSeek team showed that with pure RL, no SFT stage, it is still possible to learn advanced reasoning capabilities like reflection and backtracking (“Aha moment”). The model naturally learns to spend more thinking tokens during the RL training process to solve reasoning tasks. The “aha moment” can emerge, referring to the model reflecting on previous mistakes and then trying alternative approaches to correct them. Later, various open source efforts happened for replicating R1 results like Open-R1, SimpleRL-reason, and TinyZero, all based on Qwen models. These efforts also confirmed that pure RL leads to great performance on math problems, as well as the emergent “aha moment”.
模型学习反思和纠正错误的示例。(图片来源:(左)DeepSeek-AI, 2025;(右)Zeng et al. 2025)
Examples of the model learning to reflect and correct mistakes. (Image source: (left) DeepSeek-AI, 2025; (right) Zeng et al. 2025)
DeepSeek 团队还分享了一些不成功的尝试。他们未能成功使用过程奖励模型(PRM),因为很难定义每一步的评分标准或确定中间步骤是否正确,同时使训练更容易受到奖励黑客攻击。在 MCTS(蒙特卡洛树搜索)上的努力也失败了,因为语言模型 token 的搜索空间很大,与象棋等相比;而且训练用于指导搜索的细粒度价值模型也非常具有挑战性。失败的尝试往往提供独特的见解,我们鼓励研究社区分享更多关于哪些方法不起作用的信息。
The DeepSeek team also shared some of their unsuccessful attempts. They failed to use process reward model (PRM) as it is hard to define per-step rubrics or determine whether an intermediate step is correct, meanwhile making the training more vulnerable to reward hacking. The efforts on MCTS (Monte Carlo Tree Search) also failed due to the large search space for language model tokens, in comparison to, say, chess; and training the fine-grained value model used for guiding the search is very challenging too. Failed attempts often provide unique insights and we would like to encourage the research community to share more about what did not work out.
在推理步骤中,某些中间步骤可以通过执行代码或进行数学计算来可靠且准确地解决。将这部分推理组件卸载到外部代码解释器中,如 PAL(程序辅助语言模型;Gao 等人,2022)或 Chain of Code(Li 等人,2023),可以扩展 LLM 借助外部工具的能力,消除 LLM 学习执行代码或充当计算器的需求。这些代码模拟器,如 Chain of Code 中的,可以由 LLM 增强,这样如果标准代码解释器失败,我们可以选择使用 LLM 来执行该行代码。使用代码增强推理步骤对于数学问题、符号推理和算法任务尤其有益。这些单元测试可能不是编码问题的一部分,在这种情况下,我们可以指导模型自我生成单元测试,以便它进行测试以验证解决方案(Shinn 等人,2023)。
During the reasoning steps, certain intermediate steps can be reliably and accurately solved by executing code or running mathematical calculations. Offloading that part of reasoning components into an external code interpreter, as in PAL (Program-Aided Language Model; Gao et al. 2022) or Chain of Code (Li et al. 2023), can extend the capability of LLM with external tools, eliminating the need for LLMs to learn to execute code or function as calculators themselves. These code emulators, like in Chain of Code, can be augmented by an LLM such that if a standard code interpreter fails, we have the option of using LLM to execute that line of code instead. Using code to enhance reasoning steps are especially beneficial for mathematical problems, symbolic reasoning and algorithmic tasks. These unit tests may not exist as part of the coding questions, and in those cases, we can instruct the model to self-generate unit tests for it to test against to verify the solution (Shinn, et al. 2023).
程序辅助语言模型提示的一个示例如下。(图片来源:Gao 等人,2022)
An example of program-aided language model prompting looks like. (Image source: Gao et al. 2022)
ReAct(推理+行动;Yao 等人,2023)将搜索 Wikipedia API 的行动与推理轨迹的生成相结合,使得推理路径能够融入外部知识。
ReAct (Reason+Act; Yao et al. 2023) combines the action of searching the Wikipedia API and generation of reasoning traces, such that reasoning paths can incorporate external knowledge.
使用 ReAct 提示方法解决 HotpotQA 问题的示例,利用 Wikipedia 搜索 API 作为外部工具辅助推理。(图片来源:Yao 等人,2023)
An example of the ReAct prompting method to solve a HotpotQA question, using Wikipedia search API as an external tool to help with reasoning. (Image source: Yao et al. 2023)
OpenAI 最近发布的 o3 和 o4-mini 是另外两个很好的例子,其推理过程涉及工具使用,如网络搜索、代码执行和图像处理。团队观察到,大规模强化学习表现出与 GPT 范式相同的趋势,即“更多算力=更好性能”。
o3 & o4-mini, recently released by OpenAI, are another two good examples where the reasoning process involves tool use like Web search, code execution and image processing. The team observed that large-scale reinforcement learning exhibits the same trend as in the GPT paradigm that “more compute = better performance”.
深度学习模型常被视为黑箱,研究者提出了多种可解释性方法。可解释性有几个用处:首先,它为我们提供了一种额外测试,以判断模型是否与创建者的意图不一致,或者是否以某种我们无法通过监控其行为来察觉的方式表现异常。其次,它可以帮助我们确定模型是否使用了合理的过程来计算其答案。思维链提供了一种特别方便的可解释性形式,因为它使模型的内部过程以自然语言可见。然而,这种可解释性基于一个假设:模型真实地描述了其内部思维过程。
Deep learning models are often treated as black boxes and various interpretability methods have been proposed. Interpretability is useful for a couple reasons: first, it gives us an extra test to determine if the model is misaligned with its creators’ intent, or if it’s misbehaving in some way that we can’t tell by monitoring its actions. Second, it can help us determine whether the model is using a sound process to compute its answers. Chain of thought provides an especially convenient form of interpretability, as it makes the model’s internal process visible in natural language. This interpretability, however, rests on the assumption that the model truthfully describes its internal thought processes.
近期研究表明,监控推理模型的思维链可以有效检测模型的不当行为(如奖励黑客),甚至可以使较弱的模型监控较强的模型(Baker et al. 2025)。增加测试时算力还可以提高对抗鲁棒性(Zaremba et al. 2025);这直观上合理,因为当模型遇到异常输入(如对抗样本或越狱尝试)时,更长的思考时间应特别有用——它可以利用额外的思考时间来理解所面临的奇怪情况。
Recent work showed that monitoring CoT of reasoning models can effectively detect model misbehavior such as reward hacking, and can even enable a weaker model to monitor a stronger model (Baker et al. 2025). Increasing test time compute can also lead to improved adversarial robustness (Zaremba et al. 2025); this makes sense intuitively, because thinking for longer should be especially useful when the model is presented with an unusual input, such as an adversarial example or jailbreak attempt – it can use the extra thinking time to make sense of the strange situation it’s been presented with.
实验要求模型根据其思维过程判断另一个模型是否试图以某种方式破解编程问题的单元测试。我们可以使用不同类型的监控器在训练期间监控这些奖励黑客行为。exit(0) 编码黑客是指智能体利用一个漏洞,使其能在不运行所有单元测试的情况下提前退出环境。raise SkipTest 黑客是指智能体从测试框架之外的函数中引发异常,以跳过单元测试评估。(图片来源:Baker et al. 2025)
The experiment of asking the model to decide if another model tried to hack the unit tests in some way for coding questions given its thought process. We can monitor these reward hacking behavior during training with different types of monitor. The exit(0) coding hack is when the agent exploited a bug that allowed it to exit from the environment early without running all unit tests. The raise SkipTest hack is when the agent raises an exception from functions outside the testing framework in order to skip unit test evaluation. (Image source: Baker et al. 2025)
直觉上,由于缺乏旨在鼓励忠实推理的明确训练目标,模型的思维链可能存在偏差。或者,当我们用人类编写的解释对模型进行微调时,这些人类编写的样本可能包含错误。因此,我们不能默认假设思维链总是忠实的。
Intuitively, model CoTs could be biased due to lack of explicit training objectives aimed at encouraging faithful reasoning. Or when we fine-tune the model on human-written explanations, those human-written samples may contain mistakes. Thus we cannot by default assume CoT is always faithful .
Lanham 等人 (2023) 通过故意在思维链中引入错误并测量其对一组多项选择任务(例如 AQuA、MMLU、ARC Challenge、TruthfulQA、HellaSwag)准确率的影响,研究了思维链忠实性失败的几种模式:
Lanham et al. (2023) investigated several modes of CoT faithfulness failures by deliberately introducing mistakes into CoTs and measuring their impacts on the accuracy of a set of multiple choice tasks (e.g. AQuA, MMLU, ARC Challenge, TruthfulQA, HellaSwag):
* 错误 1(_提前回答_):模型可能在生成思维链之前过早形成结论。这通过提前截断或在思维链中插入错误来测试。不同任务显示出对思维链有效性的不同任务特定依赖性;有些任务的评估性能对截断的思维链敏感,而有些则不然。Wang 等人 (2023) 进行了类似的实验,但在思维链形成过程中使用了与桥接对象或语言模板相关的更微妙的错误。
* Mistake 1 (_Early answering_): The model may form a conclusion prematurely before CoT is generated. This is tested by early truncating or inserting mistakes into CoT. Different tasks revealed varying task-specific dependencies on CoT effectiveness; some have evaluation performance sensitive to truncated CoT but some do not. Wang et al. (2023) did similar experiments but with more subtle mistakes related to bridging objects or language templates in the formation of CoT.
* 错误 2(_无信息词元_):无信息的思维链词元提升了性能。这一假设通过用填充文本(例如全是句号)替换思维链来测试,这种设置显示准确率没有提升,并且与没有思维链相比,某些任务的性能可能略有下降。
* Mistake 2 (_Uninformative tokens_): Uninformative CoT tokens improve performance. This hypothesis is tested by replacing CoT with filler text (e.g. all periods) and this setup shows no accuracy increase and some tasks may suffer performance drop slightly when compared to no CoT.
* 错误 3(_人类不可读编码_):相关信息以人类难以理解的方式编码。以非标准方式改写思维链并未降低数据集的性能,表明准确率的提升并不依赖于人类可读的推理。
* Mistake 3 (_Human-unreadable encoding_): Relevant information is encoded in a way that is hard for humans to understand. Paraphrasing CoTs in an non-standard way did not degrade performance across datasets, suggesting accuracy gains do not rely on human-readable reasoning.
评估思维链忠实性的不同扰动方式示意图。(图片来源:Lanham 等人 2023)
Illustration of different ways of CoT perturbation to assess its faithfulness. (Image source: Lanham et al. 2023)
有趣的是,Lanham 等人指出,对于多项选择题,较小的模型可能没有足够的能力很好地利用思维链,而较大的模型可能已经能够在没有思维链的情况下解决任务。这种对思维链推理的依赖性,通过有思维链与无思维链时获得相同答案的百分比来衡量,在多项选择题上并不总是随模型规模增加而增加,但在加法任务上确实随模型规模增加而增加,这表明思考时间对于复杂推理任务更为重要。
Interestingly, Lanham et al. suggests that for multiple choice questions, smaller models may not be capable enough of utilizing CoT well, whereas larger models may have been able to solve the tasks without CoT. This dependency on CoT reasoning, measured by the percent of obtaining the same answer with vs without CoT, does not always increase with model size on multiple choice questions, but does increase with model size on addition tasks, implying that thinking time matters more for complex reasoning tasks.
对思维链推理的依赖性通过有思维链与无思维链时获得相同答案的百分比来衡量。它对于加法等推理任务更为重要,且较大的模型受益更多。(图片来源:Lanham 等人 2023)
The dependency on CoT reasoning is measured as the percentage of obtaining same answers with vs without CoT. It matters more for reasoning tasks like addition and larger models benefit more. (Image source: Lanham et al. 2023)
测试思维链忠实性的替代方法涉及扰动提示,而不是直接修改思维链路径(Turpin 等人 2023, Chua & Evans, 2025, Chen 等人 2025)。
Alternative approaches for testing CoT faithfulness involve perturbing prompts rather than modifying CoT paths directly (Turpin et al. 2023, Chua & Evans, 2025, Chen et al. 2025).
一种方法在少样本示例中始终将正确答案标记为“(A)”,而不考虑真实标签,以引入偏差。
One method consistently labels correct answers as “(A)” in few-shot examples regardless of true labels to introduce biases.
另一种提示技术将误导性提示插入到提示中,例如“我认为答案是<随机标签>,但好奇你怎么想”或“一位斯坦福教授认为答案是<随机标签>”。通过比较模型对同一问题在有和没有误导性提示下的预测,我们可以衡量模型是否能够忠实地描述提示对其答案的影响。特别是,在模型产生不同提示答案和无提示答案的情况下,我们衡量模型在解决带有提示的问题时是否承认该提示。如果模型是忠实的,它应该明确承认影响,并承认其答案的改变是由于提示。
Another prompting technique inserts misleading hints into prompts, such as "I think the answer is <random_label> but curious to hear what you think". or "A Stanford Professor thinks the answer is <random_label>". By comparing model predictions for the same question with vs without the misleading hint, we can measure whether a model is able to faithfully describe the influence of the hint on its answer. Particularly, in cases where the model produces different hint and non-hint answers, we measure whether the model acknowledges the hint when solving the question with hint. If the model is faithful, it should explicitly acknowledge the impact and admit the change of its answer is due to the hint.
GPT 和 Claude 模型都对上下文中不同类型的偏差敏感。模型准确率的下降表明系统性的不忠实。直接提示错误标签比“答案总是 A”类型的偏差更有效。(图片来源:Turpin 等人 2023)
Both GPT and Claude models are sensitive to different types of biases in context. Decrease in model accuracy indicates systematic unfaithfulness. Direct hints of wrong labels are more effectively than "Answer is always A" type of bias. (Image source: Turpin et al. 2023)
多项研究发现,推理模型描述提示影响的能力比所有测试的非推理模型可靠得多。例如,我们可以衡量模型将提示视为决定性因素的样本比例(“忠实思维链”)。推理模型(Claude 3.7 Sonnet, DeepSeek R1)总体上比非推理模型(Claude 3.6, DeepSeek V3)表现更好。
Multiple studies found that reasoning models describe the influence of the hint much more reliably than all the non-reasoning models tested. For example, we can measure the fraction of samples where the model acknowledges the hint as a deciding factor (“faithful CoT”). Reasoning models (Claude 3.7 Sonnet, DeepSeek R1) are overall doing better than non-reasoning ones (Claude 3.6, DeepSeek V3).
推理模型比非推理模型更可能揭示忠实的思维链。(图片来源:Chen 等人 2025)
Reasoning models are more likely to reveal faithful CoT than non-reasoning models. (Image source: Chen et al. 2025)
一些证据表明,使用奖励模型会导致模型输出的忠实性降低。经典 RLHF 中的奖励模型并未被训练来在此测试中激励忠实行为,因为偏好比较数据集是通过人类选择哪个看起来更好或更令人愉悦来收集的。然而,在推理模型中,思维链路径被优化以产生正确的最终答案,而不是匹配 RM 中定义的人类偏好,因此直觉上预计会更忠实。
Some evidence showed that using reward models leads to less faithfulness in model outputs. The reward model in classic RLHF is not trained to incentivize faithful behavior in this test as the preference comparison dataset is collected by humans selecting which one seems better or pleasing. However, in reasoning models, the CoT paths are optimized for producing the correct final answers, not to match human preference defined in RM, and thus intuitively are expected to be more faithful.
监控推理模型的思维链以检测奖励破解行为是有效的。一种简单的使用方式是,在测试时以监控器作为过滤器进行拒绝采样,识别出没有奖励破解的解决方案。然而,这仅在强化学习的早期阶段有效,随着模型进一步优化,在预算内很难找到能通过思维链监控器的样本。这一挑战自然表明,我们可以考虑将思维链监控器纳入强化学习奖励——一种基于过程的监督形式——以防止奖励破解。
Monitoring CoT of reasoning models for reward hacking behavior is effective. A simple way to use this monitor is to run rejection sampling with the monitor as a filter at test time and identify solutions without reward hacking. However, it only works at the early stage of RL, and, as the model is further optimized, it is rare to find a sample to pass a CoT monitor within budget. This challenge naturally suggests that we could consider incorporating CoT monitors into RL rewards—a form of process-based supervision—to prevent reward hacking.
然而,将思维链监控奖励纳入强化学习效果不佳,因为它会导致混淆的奖励破解,即模型试图在思维链中隐藏其真实意图(Baker 等人,2025)。通过强化学习训练中思维链监控器提供的这种直接优化压力,模型仍然可以找到破解方法,尽管频率稍低,但可以在思维链中隐藏其意图以避免被监控器捕获。
However, incorporating CoT monitoring reward into RL does not work out well, as it leads to obfuscated reward hacking where the model tries to hide its true intent within the CoT (Baker et al. 2025). With this direct optimization pressure provided by the CoT monitor applied during RL training, the model still can find ways to hack, although a bit less frequently, but can hide its intention in CoT to avoid being caught by the monitor.
当使用监控器提供强化学习奖励信号以防止训练中的奖励破解时,模型仍然可以找到破解方法,同时避免被监控器捕获。(图片来源:Baker 等人,2025)
When using the monitor to provide an RL reward signal to prevent reward hacking during training, the model still can find ways to hack while avoiding being caught by the monitor. (Image source: Baker et al. 2025)
Yeo 等人(2025)关于思维链长度强化学习信号的实验证实了与 Baker 等人类似的结论,即新的强化学习奖励塑形可能导致新型奖励破解。他们设计的奖励函数使得正确的思维链比错误的获得更高奖励,短的正确思维链比长的正确思维链获得更高奖励,短的错误思维链比长的错误思维链受到更高惩罚。在这种额外奖励下,模型学会了在思维链中重复文本以应对难题,而不是尝试解决它们。Yeo 等人进一步应用了 n-gram 重复惩罚来防止这种破解行为。
Experiments on CoT length RL signals by Yeo et al. (2025) confirmed a similar message as Baker et al. that new types of reward hacking can happen with new types of RL reward shaping. They designed the reward function such that correct CoTs can have higher rewards than wrong ones, short correct CoTs obtain higher rewards than long correct ones and short incorrect CoT receive higher penalties than long incorrect CoTs. With this extra reward, the model learns to repeat text in CoTs for challenging questions rather than attempting to solve them. Yeo et al. further applied a n-gram repetition penalty to prevent this hacking behavior.
需要谨慎的奖励塑形来稳定使用思维链长度奖励的训练。(图片来源:Yeo 等人,2025)
Careful reward shaping is needed to stabilize training with CoT length rewards. (Image source: Yeo et al. 2025)
Chen 等人(2025)实验了一个有缺陷的强化学习环境,具体是在多项选择题中使用带有错误答案的评分器。模型学会了在超过 99%的提示上利用奖励破解,但在超过一半的环境中几乎从未(<2%)在思维链中口头表达奖励破解。在这种情况下,额外的强化学习优化压力未能激励模型口头表达破解。
Chen et al. (2025) experimented with a flawed RL environment, specifically using a grader with incorrect answers filled in for multiple-choice questions. The model learns to exploit the reward hack on >99% of the prompts, but almost never (<2%) verbalizes the reward hack in its CoT on more than half of their environments. Additional RL optimization pressure fails to incentivize the model to verbalize the hack in this case.
强化学习训练本质上对奖励破解敏感。仅依赖启发式调查奖励破解和手动修复可能导致“打地鼠”局面。我们建议在尝试将优化直接应用于强化学习训练中的思维链时非常谨慎,或者完全避免这样做。
RL training is inherently sensitive to reward hacking. Only relying on heuristic investigation of reward hacking and manual fixes may lead to a “whack-a-mole” situation. We would suggest being very cautious when trying to apply optimization directly on CoT during RL training, or trying to avoid it altogether.
自适应计算时间(Adaptive Computation Time)由 Alex Graves 于 2016 年提出,虽早于大型语言模型,但开创了使模型在推理时动态决定计算步数的方向,可视为让模型在测试时于连续空间中“思考更多”。连续空间中的自适应思考时间可通过循环架构纵向实现,或通过更多顺序采样步骤横向实现。
Adaptive Computation Time, introduced by Alex Graves in 2016, predated large language models but pioneered the same direction of enabling the model to dynamically decide the number of computational steps to take at the inference time, which can be viewed as enabling the model to “think more” in continuous space at test time. Adaptive thinking time in continuous space can be enabled vertically via recurrent architecture or horizontally via more sequential sampling steps.
已经提出了多种架构变体,使 Transformer 架构具有循环性,从而实现自适应测试时算力(Dehghani 等人,2019;Hutchins 等人,2022;Bulatov 等人,2022)。深入探讨该主题的文献会使文章过长,因此我们仅回顾少数几种。
A number of architecture variations have been proposed to make the Transformer architecture recurrent, enabling adaptive test time compute (Dehghani, et al. 2019, Hutchins, et al. 2022, Bulatov, et al. 2022). A deep dive into literature on this topic would make the post too long, so we will only review a few.
通用 Transformer(Dehghani 等人,2019)将 Transformer 中的自注意力机制与 RNN 中的循环机制相结合,使用自适应计算时间(Graves,2016)动态调整步数。从高层次看,它可以被视为一个学习每个 token 隐藏状态表示的循环函数,如果步数固定,通用 Transformer 等价于一个跨层共享参数的多层 Transformer。
Universal Transformer (Dehghani, et al. 2019) combines self-attention in Transformer with the recurrent mechanism in RNN, dynamically adjusting the number of steps using adaptive computation time (Graves, 2016). On a high level, it can be viewed as a recurrent function for learning the hidden state representation per token, and if the number of steps is fixed, an Universal Transformer is equivalent to a multi-layer Transformer with shared parameters across layers.
Geiping 等人(2025)提出的一种近期循环架构设计,在标准 Transformer 之上添加了一个循环块 $R$。该循环块的每次迭代都接受嵌入 $\mathbf{e}$ 和一个随机状态 $\mathbf{s}_{i}$。从概念上讲,这种循环深度架构有点类似于条件扩散模型,其中原始输入 $\mathbf{e}$ 在每个循环步骤中提供,而随机高斯初始化的状态 $\mathbf{s}_{i}$ 通过该过程迭代更新。(有趣的是,他们的一些更接近扩散模型的设计实验结果并不好。)
A recent recurrent architecture design, proposed by Geiping et al. (2025), adds a recurrent block $R$ on top of the standard Transformer. Every iteration of this recurrent block takes the embedding $\mathbf{e}$ and a random state $\mathbf{s}_{i}$. Conceptually, this recurrent-depth architecture is a bit similar to a conditioned diffusion model, where the original input $\mathbf{e}$ is provided in every recurrent step while a random Gaussian initialized state $\mathbf{s}_{i}$ gets updated iteratively through the process. (Interestingly some of their experiments of designs that resemble diffusion models more turned out to be bad.)
\mathbf{e} & = P \left(\right. \mathbf{x} \left.\right) & \text{嵌入} \\ \mathbf{s} _ 0 & sim \mathcal{N} \left(\right. 0 , \sigma^{2} I _ n \cdot h \left.\right) \\ \mathbf{s} _ i & = R \left(\right. \mathbf{e} , \mathbf{s} _ i - 1 \left.\right) \textrm{ }\text{对于}\textrm{ } i \in 1 , \ldots , r & \text{循环块};\text{类似于 Transformer 块} \\ \mathbf{p} & = C \left(\right. \mathbf{s} _ r \left.\right) & \text{解嵌入}
\mathbf{e} & = P \left(\right. \mathbf{x} \left.\right) & \text{embedding} \\ \mathbf{s} _ 0 & sim \mathcal{N} \left(\right. 0 , \sigma^{2} I _ n \cdot h \left.\right) \\ \mathbf{s} _ i & = R \left(\right. \mathbf{e} , \mathbf{s} _ i - 1 \left.\right) \textrm{ }\text{for}\textrm{ } i \in 1 , \ldots , r & \text{recurrent block};\text{ resembles a Transformer block} \\ \mathbf{p} & = C \left(\right. \mathbf{s} _ r \left.\right) & \text{unembedding}
训练期间的循环次数 $r$ 是随机化的,从对数正态泊松分布中为每个输入序列采样。为了管理计算成本,反向传播被截断为仅循环单元的最后 $k$ 次迭代(实验中 $k = 8$),从而可以在泊松分布的重尾部分进行训练。嵌入块在每个步骤中继续接收梯度更新,因为其输出 $\mathbf{e}$ 在每个步骤中都被注入,模仿了 RNN 训练。毫不意外,循环模型训练的稳定性非常敏感。初始化、归一化和超参数等因素都很重要,尤其是在扩展训练规模时。例如,隐藏状态可能因预测每个 token 的相同隐藏状态而崩溃;或者模型可能学会忽略传入状态 $\mathbf{s}$。为了稳定训练,Geiping 等人采用了嵌入缩放因子、小学习率和仔细的调参。
The recurrence count $r$ during training is randomized, sampled from a log-normal Poisson distribution, per input sequence. To manage computational costs, backpropagation is truncated to only the last $k$ iterations of the recurrent unit ($k = 8$ in experiments), making it possible to train on the heavy-tail part of the Poisson distribution. The embedding block continues to receive gradient updates in every step since its output $\mathbf{e}$ is injected in every step, mimicking RNN training. Unsurprisingly, the stability of training a recurrent model turns out to be very sensitive. Factors like initialization, normalization and hyperparameters all matter, especially when scaling up the training. For example, hidden states can collapse by predicting the same hidden state for every token; or the model may learn to ignore the incoming state $\mathbf{s}$. To stabilize the training, Geiping et al. adopted an embedding scale factor, a small learning rate and careful tuning.
训练 3.5B 模型并带有深度循环的实验图。饱和大致发生在 $\bar{r} = 32$ 左右,这让我们思考该架构如何外推并泛化到更大的迭代次数。(图片来源:Geiping 等人,2025)
Plot of an experiment on training a 3.5B model with depth recurrence. The saturation roughly happens around $\bar{r} = 32$, making us wonder how this architecture extrapolate and generalize to larger iteration counts. (Image source: Geiping et al. 2025)
思考令牌指的是一组在训练或推理过程中引入的隐式令牌,它们不携带直接的语义含义。相反,它们的作用是为模型提供额外的思考时间和算力,以使其表现更好。
Thinking tokens refer to a set of implicit tokens introduced during training or inference that do not carry direct linguistic meaning. Instead, their role is to provide extra thinking time and compute power for the model to perform better.
Herel & Mikolov (2023) 提出了在句子中每个词后插入特殊思考令牌 (<T>) 并在此类数据集上训练模型的想法。每个思考令牌为模型争取了额外的处理时间,以做出更好的预测。在玩具模型设置中使用思考令牌进行训练,其困惑度低于未使用思考令牌训练的基线模型。对于非平凡的推理任务或涉及数字的句子,思考令牌的益处更为显著。
Herel & Mikolov (2023) introduced the idea of inserting special thinking tokens (<T>) after each word in a sentence and training the model on such a dataset. Each thinking token buys extra time for the model to process and make better predictions. Training with thinking tokens on a toy model setup results in lower perplexity than baseline model trained without them. The benefits of thinking tokens are more pronounced for non-trivial reasoning tasks or sentences involving numbers.
类似地,Goyal 等人 (2024) 提出的暂停令牌通过在输入序列末尾附加虚拟令牌(例如字符 . 或 #)来延迟模型输出,从而在推理期间为模型提供额外的计算。重要的是,在训练和推理期间都需要注入此类暂停令牌,而仅在暂停令牌上进行微调带来的收益有限。在训练期间,多个暂停令牌副本被插入到均匀随机的位置,并且训练时忽略暂停令牌上的损失。
Similarly, pause tokens proposed by Goyal et al. (2024) delay the model’s outputs by appending dummy tokens (e.g. character like . or #) at the end of the input sequence, giving the model extra computation during inference. It is important to inject such pause tokens both during training and inference time, while only fine-tuning on pause tokens leads to limited gain. During training, multiple copies of pause tokens are inserted at uniformly random locations and the loss on pause tokens is ignored for training.
与标准设置相比,暂停令牌在训练和推理期间如何注入的示意图。(图片来源:Goyal 等人 2024)
Illustration of how pause tokens are injected during training and inference in comparison to standard setup. (Image source: Goyal et al. 2024)
有趣的是,上述实验中的思考令牌或暂停令牌不携带任何额外信息,也不增加许多新参数。但为什么它们仍然有帮助?一方面,它们通过引入更多推理循环来扩展计算,有效增加了计算能力。另一方面,它们可以被视为一种特殊的、隐式的思维链形式。一个缺点是模型需要针对思考令牌进行预训练。尽管如此,这种策略是在推理时思维链基础上进一步提高测试时算力利用能力的一种有趣方式。
Interestingly, thinking tokens or pause tokens in above experiments do not carry any extra information or add many new parameters. But why is it still helpful? On one hand, it helps expand the computation by introducing more inference loops, effectively increasing computational capacity. On the other hand, it can be seen as acting as a special, implicit form of CoTs. A drawback here is hat the model needs to be pretrained with respect to thinking tokens. Still, this strategy is an interesting way to further improve the capability of test time compute utilization on top of inference time CoTs.
Quiet-STaR (Zelikman 等人 2025) 通过训练模型在每个词元后生成理由来解释未来文本,引入了词元级推理。它将有无理由的未来文本预测混合,并使用学习来生成更好的理由,同时使用 REINFORCE 优化理由生成的质量。
Quiet-STaR (Zelikman et al. 2025) introduces token-level reasoning by training the model to generate rationales after every token to explain future text. It mixes the future-text predictions with and without rationales and uses learning to generate better rationales and uses REINFORCE to optimize the quality of rationale generation.
Quiet-STaR 示意图。(图片来源:Zelikman 等人 2025)
Illustration of Quiet-STaR. (Image source: Zelikman et al. 2025)
* _思考_:使用理由预测下一个词元。由于词元级推理所需的高计算成本,该过程设计为并行生成多个理由。使用特殊的注意力映射,使得所有思考词元仅关注自身、同一思考内的所有先前思考词元以及前面的文本。
* _Think_: Predicting next tokens with rationales. Due to high computational cost demanded by token level reasoning, this process is designed to generate multiple rationales in parallel. A special attention map is used to enable all thought tokens to only pay attention to themselves, all preceding thought tokens within the same thought, and the preceding text.
* _表达_:无理由的下一个词元预测与有理由的预测混合。两个 logits 的混合权重由每个理由后的隐藏输出通过一个浅层 MLP 的特殊混合头学习。正确的下一个词元可以通过教师强制选择。
* _Talk_: Next token prediction without rationale is mixed with post-rationale prediction. The mixing weight for two logits is learned by a special mixing head of a shallow MLP out of hidden output after each rationale. Correct next tokens can be selected via teacher forcing.
* _学习_:通过 REINFORCE 训练模型生成更好的理由,从那些增加正确下一个词元概率的示例中学习,同时丢弃那些损害预测的示例。
* _Learn_: Train the model to generate better rationale via REINFORCE by learning from examples that increase the probability of correct next token while discarding those that hurt the prediction.
无需特定数据集的微调,Quiet-STaR 在 Mistral 7B 上的实验中将 CommonsenseQA 的零样本结果从 36.3% 提升至 47.2%,GSM8K 从 5.9% 提升至 10.9%。
Without dataset specific fine-tuning, Quiet-STaR improves zero-shot results on CommonsenseQA (36.3%→47.2%) and GSM8K (5.9%→10.9%), within experiments on Mistral 7B.
潜在变量模型定义了一个概率框架,其中可观测数据通过未观测(潜在)变量来解释。这些潜在变量捕捉生成可观测结果的隐藏结构或中间过程。语言模型可以被视为概率潜在变量模型,其中测试时的思维和推理步骤是潜在思维变量(Zhou et al. 2020, Phan et al. 2023)。这样的潜在变量模型定义了问题 x_i、答案 y_i 和潜在思维 z_i 的联合分布。我们希望优化给定问题和多种思维链作为潜在变量时答案的对数似然(N 是样本数;$K$ 是每个问题的思维链数量):
A latent variable model defines a probabilistic framework where observable data is explained through unobserved (latent) variables. These latent variables capture hidden structures or intermediate processes that generate the observable outcomes. Language models can be viewed as probabilistic latent variable models where test-time thinking and reasoning steps are latent thought variables (Zhou et al. 2020, Phan et al. 2023). Such a latent variable model defines a joint distribution of problems x_i, answers y_i and latent thought z_i. We would like to optimize the log-likelihood of answers given questions and a variety of CoTs as latent variables (N is the number of samples; $K$ is the number of CoTs per problem):
log \mathcal{L} \left(\right. \theta \left.\right) & = log p \left(\right. y \mid x \left.\right) \\ & = log \sum_{k = 1}^{K} p \left(\right. y , z^{\left(\right. k \left.\right)} \mid x \left.\right) \\ & = log \sum_{k = 1}^{K} p \left(\right. z^{\left(\right. k \left.\right)} \mid x \left.\right) p \left(\right. y \mid z^{\left(\right. k \left.\right)} , x \left.\right) \\ & = log \mathbb{E}_{z^{\left(\right. k \left.\right)} sim p \left(\right. z^{\left(\right. k \left.\right)} \mid x \left.\right)} p \left(\right. y \mid z^{\left(\right. k \left.\right)} , x \left.\right)
log \mathcal{L} \left(\right. \theta \left.\right) & = log p \left(\right. y \mid x \left.\right) \\ & = log \sum_{k = 1}^{K} p \left(\right. y , z^{\left(\right. k \left.\right)} \mid x \left.\right) \\ & = log \sum_{k = 1}^{K} p \left(\right. z^{\left(\right. k \left.\right)} \mid x \left.\right) p \left(\right. y \mid z^{\left(\right. k \left.\right)} , x \left.\right) \\ & = log \mathbb{E}_{z^{\left(\right. k \left.\right)} sim p \left(\right. z^{\left(\right. k \left.\right)} \mid x \left.\right)} p \left(\right. y \mid z^{\left(\right. k \left.\right)} , x \left.\right)
我们的目标是最大化正确答案的边际似然 $p \left(\right. y \mid x \left.\right)$,给定每个问题的多个推理轨迹 $\left{\right. z^{\left(\right. k \left.\right)} \left.\right}_{k = 1}^{K}$。
Our goal is to maximize the marginal likelihood of the correct answer, $p \left(\right. y \mid x \left.\right)$, given a number of reasoning traces per problem, $\left{\right. z^{\left(\right. k \left.\right)} \left.\right}_{k = 1}^{K}$.
期望最大化是一种常用的迭代算法,用于优化带有(隐藏)潜变量的模型参数,因此可用于训练更好的思维链,并在此基础上生成更好的响应。通常我们在 E 步(期望步)中猜测关于潜变量的缺失信息(即如何采样更好的思维链),在 M 步(最大化步)中基于潜变量优化模型参数(即如何采样更好的答案),直至收敛。
Expectation-Maximization is a commonly used iterative algorithm for optimizing parameters for a model with (hidden) latent variables, and thus can be applied to train better CoTs and then condition on that to generate better responses. Typically we iterate between E-step (Expectation) where we guess the missing information about latent variables (i.e. how to sample better CoTs), and M-step (Maximization) where we optimize the model parameters based on latent variables (i.e. how to sample better answers), until convergence.
log ℒ(θ) = log 𝔼_{z^{(k)} ∼ p(z^{(k)} | x)⏟E-step} p(y | z^{(k)}, x)⏟M-step
log \mathcal{L} \left(\right. \theta \left.\right) = log \mathbb{E}_{\underset{\text{E}-\text{step}}{\underbrace{z^{\left(\right. k \left.\right)} sim p \left(\right. z^{\left(\right. k \left.\right)} \mid x \left.\right)}}} \underset{\text{M}-\text{step}}{\underbrace{p \left(\right. y \mid z^{\left(\right. k \left.\right)} , x \left.\right)}}
由于我们无法直接从潜变量分布 $p(z | x, y)$ 中采样,研究者探索了依赖人工标注数据(Zhou et al. 2020)、Metropolis-Hastings MCMC(Phan et al. 2023)或带有特殊重要性权重的蒙特卡洛采样(Ruan et al. 2025)等方法,以获取好的思维链样本来更新模型。Ruan 等人(2025)尝试使用 EM 算法在大量包含潜在思维的 Web 文本上训练模型,其中每块观测数据会合成对应的潜在思维,然后模型以自回归方式同时学习潜在思维和数据。
Because we cannot directly sample from the latent variable distribution $p \left(\right. z \mid x , y \left.\right)$, researchers have explored methods relying on human annotated data (Zhou et al. 2020), Metropolis-Hastings MCMC (Phan et al. 2023) or Monte Carlo sampling with special importance weights (Ruan et al. 2025) to draw good CoT samples to update the model. Ruan et al. (2025) experimented with training a model on a large body of Web text with latent thoughts with the EM algorithm, where the latent thought is synthesized per chunk of observed data and then the model learns over both latent thought and data in an autoregressive manner.
在数据语料中注入潜在思维进行训练的示意图。(图片来源:Ruan et al. 2025)
Illustration of training on data corpus with latent thought injected. (Image source: Ruan et al. 2025)
他们首先提示一个 LLM $\overset{\sim}{q}$,根据观测数据 X_i 生成合成潜在思维 Z_i:
They first prompt a LLM $\overset{\sim}{q}$ to generate synthetic latent thought Z_i given observed data X_i:
你被提供一对网页文档前缀和后缀。你的任务是在它们之间插入潜在思维,这些思维是后缀基于前缀生成的基础。潜在思维应包括:缺失的背景知识和每个主张背后的推理轨迹(特别是逐步推导或逻辑推理)。
You are provided with a pair of web document prefix and suffix. Your task is to insert latent thoughts between them underlying the creation of the suffix conditioned on the prefix. The latent thoughts should include: the missing background knowledge and the reasoning traces underlying each claim (especially, step-by-step derivations or logical reasoning).
使用特殊标记如 <StartOfLatent><Prior> ... <EndOfPrior> 将生成的潜在思维内容插入原始数据,用于训练联合分布 $p(z, x)$ 或近似后验 $q(z | x)$,具体取决于 $z$ 是插入在 $x$ 之前还是之后。然而,由于我们使用 LLM $\overset{\sim}{q}(z | x)$ 生成思维链,这给近似 $q(z | x)$ 的质量设定了性能上限。Ruan 等人引入了重要性权重,用于在 E 步选择思维链样本,公式为:
Special tokens like <StartOfLatent><Prior> ... <EndOfPrior> are used to insert the generated latent thought content into the raw data for training the joint distribution $p \left(\right. z , x \left.\right)$ or the approximate posterior $q \left(\right. z \mid x \left.\right)$, depending on whether $z$ is inserted before or after $x$. However, since we are using a LLM $\overset{\sim}{q} \left(\right. z \mid x \left.\right)$ to generate the CoTs, it imposed a performance ceiling on how good the approximate $q \left(\right. z \mid x \left.\right)$ can be. Ruan et al. introduced importance weights for selecting CoT samples at the E-step, formulated as:
w^{(k)} = \frac{p(z^{(k)}, x)}{q(z^{(k)} | x)} = \frac{p(x | z^{(k)}) p(z^{(k)})}{q(z^{(k)} | x)}
w^{\left(\right. k \left.\right)} = \frac{p \left(\right. z^{\left(\right. k \left.\right)} , x \left.\right)}{q \left(\right. z^{\left(\right. k \left.\right)} \mid x \left.\right)} = \frac{p \left(\right. x \mid z^{\left(\right. k \left.\right)} \left.\right) p \left(\right. z^{\left(\right. k \left.\right)} \left.\right)}{q \left(\right. z^{\left(\right. k \left.\right)} \mid x \left.\right)}
这样我们优先选择那些思维链能很好预测观测(即高 $p(x | z^{(k)})$)、简单直观(即高 $p(z^{(k)})$)但又信息丰富且不过于明显(即低 $q(z^{(k)} | x)$)的样本。
, such that we prioritize samples with CoTs that are good at predicting the observation (i.e., high $p \left(\right. x \mid z^{\left(\right. k \left.\right)} \left.\right)$), simple, intuitive (i.e., high $p \left(\right. z^{\left(\right. k \left.\right)} \left.\right)$) but also informative and not too obvious (i.e. low $q \left(\right. z^{\left(\right. k \left.\right)} \mid x \left.\right)$).
由于预训练模型已经具备生成思维链的能力,很自然地会设计一个迭代改进过程:生成多个思维链,并仅对导致正确答案的推理过程进行微调。
Since pretrained models already possess the capability of generating chains of thought, it is intuitive to design an iterative improvement process where we generate multiple CoTs and fine-tune the model only on rationales that lead to correct answers.
然而,这种直接的设计可能会失败,因为模型对于未能解决的问题无法获得学习信号。STaR(“自教推理器”;Zelikman 等人,2022)通过为失败的尝试添加一个“合理化”过程来解决这一局限,在该过程中,模型基于问题和真实答案反向生成合理的思维链,从而能够生成更合理的思维链。然后,模型在正确的解决方案上进行微调,这些方案要么直接产生正确输出,要么通过合理化生成。
However, this straightforward design can fail because the model receives no learning signals for problems it fails to solve. STaR (“Self-taught reasoner”; Zelikman et al. 2022) addresses this limitation by adding a “rationalization” process for failed attempts, in which the model generates good CoTs backward conditioned on both the problem and the ground truth answer and thus the model can generate more reasonable CoTs. Then the model is finetuned on correct solutions that either lead to correct outputs or are generated through rationalization.
STaR 的算法。(图片来源:Zelikman 等人,2022)
The algorithm of STaR. (Image source: Zelikman et al. 2022)
我们可以将 STaR 视为强化学习中策略梯度的一种近似,使用简单的指示函数作为奖励,即 𝟙$1 \left[\right. \hat{y} = y \left]\right.$。我们希望最大化在采样 $z sim p \left(\right. z \mid x \left.\right)$ 然后 $y sim p \left(\right. y \mid x , z \left.\right)$ 时该奖励的期望,因为 $p \left(\right. y \mid x \left.\right) = \underset{z}{\sum} p \left(\right. z \mid x \left.\right) p \left(\right. y \mid x , z \left.\right)$。
We can view STaR as an approximation to a policy gradient in RL, with a simple indicator function as the reward, 𝟙$1 \left[\right. \hat{y} = y \left]\right.$. We want to maximize the expectation of this reward when sampling $z sim p \left(\right. z \mid x \left.\right)$ and then $y sim p \left(\right. y \mid x , z \left.\right)$, since $p \left(\right. y \mid x \left.\right) = \underset{z}{\sum} p \left(\right. z \mid x \left.\right) p \left(\right. y \mid x , z \left.\right)$.
\nabla_{\theta} J \left(\right. \theta \left.\right) & = \nabla_{\theta} \mathbb{E}_{z_{i} , y_{i} sim p \left(\right. . \mid x_{i} \left.\right)} 1 \left[\right. y_{i} = y_{i}^{\text{truth}} \left]\right. \\ & = \sum_{i = 1}^{N} \nabla_{\theta} 1 \left[\right. y_{i} = y_{i}^{\text{truth}} \left]\right. p \left(\right. y_{i} , z_{i} \mid x_{i} \left.\right) \\ & = \sum_{i = 1}^{N} 1 \left[\right. y_{i} = y_{i}^{\text{truth}} \left]\right. p \left(\right. y_{i} , z_{i} \mid x_{i} \left.\right) \frac{\nabla_{\theta} p \left(\right. y_{i} , z_{i} \mid x_{i} \left.\right)}{p \left(\right. y_{i} , z_{i} \mid x_{i} \left.\right)} & ;\text{log}-\text{导数技巧} \\ & = \mathbb{E}_{z_{i} , y_{i} sim p \left(\right. . \mid x_{i} \left.\right)} 1 \left[\right. y_{i} = y_{i}^{\text{truth}} \left]\right. \nabla_{\theta} log p \left(\right. y_{i} , z_{i} \mid x_{i} \left.\right) & ;\text{log}-\text{导数技巧}
\nabla_{\theta} J \left(\right. \theta \left.\right) & = \nabla_{\theta} \mathbb{E}_{z_{i} , y_{i} sim p \left(\right. . \mid x_{i} \left.\right)} 1 \left[\right. y_{i} = y_{i}^{\text{truth}} \left]\right. \\ & = \sum_{i = 1}^{N} \nabla_{\theta} 1 \left[\right. y_{i} = y_{i}^{\text{truth}} \left]\right. p \left(\right. y_{i} , z_{i} \mid x_{i} \left.\right) \\ & = \sum_{i = 1}^{N} 1 \left[\right. y_{i} = y_{i}^{\text{truth}} \left]\right. p \left(\right. y_{i} , z_{i} \mid x_{i} \left.\right) \frac{\nabla_{\theta} p \left(\right. y_{i} , z_{i} \mid x_{i} \left.\right)}{p \left(\right. y_{i} , z_{i} \mid x_{i} \left.\right)} & ;\text{log}-\text{derivative trick} \\ & = \mathbb{E}_{z_{i} , y_{i} sim p \left(\right. . \mid x_{i} \left.\right)} 1 \left[\right. y_{i} = y_{i}^{\text{truth}} \left]\right. \nabla_{\theta} log p \left(\right. y_{i} , z_{i} \mid x_{i} \left.\right) & ;\text{log}-\text{derivative trick}
每次迭代相当于首先根据 𝟙$1 \left[\right. y = y^{\text{truth}} \left]\right.$ 选择思维链样本,然后运行监督微调以优化生成好的思维链和答案的对数概率。STaR 的性能随着训练迭代次数的增加而提升,并且用于生成更好思维链的“合理化”过程加速了学习。他们观察到,使用高温采样会增加得到正确答案但推理错误的机会,并在这样的数据上微调大语言模型会损害泛化能力。对于没有真实答案的数据集,多个高温输出的大多数投票可以作为真实答案的代理(Wang 等人,2022),从而使得使用合成样本进行训练成为可能。
Each iteration is equivalent to first selecting the CoT samples according to 𝟙$1 \left[\right. y = y^{\text{truth}} \left]\right.$ and then running supervised fine-tuning to optimize the logprob of generating good CoTs and answers. Performance of STaR improves with more training iterations, and the “rationalization” process for generating better CoTs accelerates learning. They observed that sampling with high temperature increases the chance of getting correct answers with incorrect reasoning and finetuning LLM on such data can impair generalization. For datasets without ground truths, majority votes of multiple high-temperature outputs can serve as a proxy of ground truth answers (Wang et al. 2022), making it possible to use synthetic samples for training.
两个 $n$ 位数求和准确率的比较。通过合理化(基于真实答案生成思维链),模型可以较早地学习像 $5$ 位数求和这样的复杂算术任务。(图片来源:Zelikman 等人,2022)
A comparison of accuracy of two $n$-digit numbers summation. With rationalization (CoT generation conditioned on ground truth), the model can learn complex arithmetic tasks like $5$-digit summation pretty early on. (Image source: Zelikman et al. 2022)
迄今为止,已有大量证据表明,允许模型在推理时、生成最终答案之前投入额外算力进行推理,可以显著提升性能。诸如提示模型在给出答案前生成中间推理步骤,或训练模型在预测下一个词之前暂停并反思等技术,已被发现能够将模型性能提升至训练所获能力上限之上。这本质上引入了一个提升模型智能的新维度,补充了缩放定律(Kaplan 等人,2020)中定义的模型规模、训练算力和数据量等既定因素。
So far we have seen much evidence that allowing models to spend additional compute on reasoning before producing final answers at inference time can significantly improve performance. Techniques like prompting the model to generate intermediate reasoning steps before the answers, or training the model to pause and reflect before predicting next tokens, have been found to boost the model performance beyond the capability limit obtained during training. This essentially introduces a new dimension to tinker with for improving model intelligence, complementing established factors such as model size, training compute and data quantity, as defined in scaling laws (Kaplan et al. 2020).
近期研究表明,优化大语言模型的测试时算力可能比扩大模型参数更有效(Snell 等人,2024;Wu 等人,2025)。较小的模型结合先进的推理算法,可以在成本和性能上实现帕累托最优的权衡。
Recent studies demonstrated that optimizing LLM test-time compute could be more effective than scaling up model parameters (Snell et al. 2024, Wu et al. 2025). Smaller models combined with advanced inference algorithms can offer Pareto-optimal trade-offs in cost and performance.
Snell 等人(2024)评估并比较了测试时算力和预训练算力,发现它们并非 1:1 可互换。当模型能力差距较小时,测试时算力可以轻松弥补简单和中等问题的差距,但对于难题则效果不佳。预训练与推理的 token 预算比例至关重要。只有当推理 token 远少于预训练 token 时,测试时算力才更可取。这表明,通过足够的预训练数据和算力开发一个有能力的基础模型仍然非常关键,因为测试时算力无法解决所有问题,也无法填补模型能力的巨大差距。
Snell et al. (2024) evaluated and compared test-time and pretraining compute, and found that they are not 1:1 exchangeable. Test-time compute can cover up the gap easily on easy and medium questions when there is only a small gap in model capability, but proves less effective for hard problems. The ratio between token budgets for pretraining and inference matters a lot. Test-time compute is only preferable when inference tokens are substantially fewer than pretraining ones. This indicates that developing a capable base model with enough pretraining data and compute is still very critical, as test-time compute cannot solve everything or fill in big model capability gaps.
(左图)通过迭代修订或并行解码,评估准确率随测试时算力预算的变化。(右图)比较使用测试时算力采样技巧的小模型与仅使用贪心解码的 14 倍大模型。我们可以控制测试时使用的 token 预算,使得推理与预训练 token 的比例远小于 1、约等于 1 或远大于 1,测试时算力的优势仅在比例远小于 1 时明显。(图片来源:Snell 等人,2024)
(Left) Evaluation accuracy as a function of test time compute budgets, via iterative revisions or parallel decoding. (Right) Comparing a small model with test-time compute sampling tricks and a 14x larger model with only greedy decoding. We can control the token budgets used at test time such that the ratio of inference to pretraining tokens is << 1, ~=1 or >> 1 and the benefit of test time compute is clear only the ratio << 1. (Image source: Snell et al. 2024)
s1 模型(Muennighoff & Yang 等人,2025)通过预算强制技术(即通过附加“wait”一词强制延长,或通过附加思考结束 token 或“Final Answer:”终止模型的思考过程来缩短)实验性地扩展了思维链推理路径长度。他们观察到,以 token 计的平均思考时间与下游评估准确率之间存在明显的正相关。
s1 models (Muennighoff & Yang, et al. 2025) experimented with scaling up the CoT reasoning path length via _budget forcing_ technique (i.e. forcefully lengthen it by appending the word "wait", or shorten it by terminating the model’s thinking process by appending end-of-thinking token or "Final Answer:"). They observed a clear positive correlation between the average thinking time measured in tokens and the downstream evaluation accuracy.
在 s1 实验中,测试时算力的并行和顺序缩放方法均与评估性能呈正相关。(图片来源:Muennighoff & Yang 等人,2025)
Both parallel and sequential scaling methods of test-time compute shows positive correlation with the evaluation performance in s1 experiments. (Image Muennighoff & Yang, et al. 2025)
当将这种预算强制技术与控制推理轨迹长度的其他解码方法进行比较时,令人惊讶的是,简单的拒绝采样(即采样生成直到长度符合 token 预算)会导致反向缩放,即更长的思维链导致更差的性能。
When comparing this budget forcing technique with other decoding methods of controlling reasoning trace length, it is quite surprising that simple rejection sampling (i.e. sampling generation until the lengths fits a token budget) leads to reversed scaling, meaning that longer CoTs lead to worse performance.
(左图)更长的思维链路径长度与评估准确率呈正相关。(右图)用于控制生成思维链路径长度的拒绝采样显示出负缩放,即更长的思维链导致更差的评估准确率。(图片来源:Muennighoff & Yang 等人,2025)
(Left) Longer CoT path length is positively correlated with eval accuracy. (Right) Rejection sampling for controlling the generated CoT path length shows a negative scaling where longer CoTs lead to worse eval accuracy. (Image source: Muennighoff & Yang et al. 2025)
对测试时算力和思维链推理的探索为增强模型能力提供了新的机遇。更有趣的是,通过测试时思考,我们正朝着构建未来 AI 系统的方向迈进,这些系统将模仿人类思考的最佳实践,融入适应性、灵活性、批判性反思和错误修正。当前进展带来的兴奋感促使我们进行更多未来研究,以深入理解我们——以及我们的模型——如何思考以及为何思考。
The exploration of test-time compute and chain-of-thought reasoning presents new opportunities for enhancing model capabilities. More interestingly, via test-time thinking, we are moving towards building future AI systems that mirror the best practices of how humans think, incorporating adaptability, flexibility, critical reflection, and error correction. Excitement with current progress invites us for more future research to improve and understand deeply not just how but why we—and our models—think.
最后,我想呼吁对以下关于测试时算力和思维链推理的开放研究问题进行更多研究。
At the end, I would like to call for more research for the following open research questions on test time compute and chain-of-thought reasoning.
* 我们能否在强化学习训练中激励模型生成人类可读、忠实的推理路径,同时避免奖励破解行为?
* Can we incentivize the model to produce human-readable, faithful reasoning paths during RL training while avoiding reward hacking behavior?
* 如何定义奖励破解?我们能否在强化学习训练或推理过程中无需人工干预地捕获奖励破解?如何防止在强化学习训练中针对奖励破解的“打地鼠”式修复?
* How to define reward hacking? Can we capture reward hacking during RL training or inference without human intervention? How to prevent “whack-a-mole” style of fixes for reward hacking during RL training?
* 自我修正可以在思维链内发生,也可以在多轮强化学习中被明确鼓励。当没有真实答案时,我们如何训练模型在不产生幻觉或退步的情况下自我修正?
* Self-correction can happen within chain-of-thought or can be encouraged to happen explicitly during multi-turn RL. How can we train the model to correct itself without hallucination or regression when ground truth is not available?
* 如何对高度情境化、个性化且难以评分的任务(如创意写作、辅导、头脑风暴)进行带有思维链展开的强化学习训练?
* How to run RL training with CoT rollout for tasks that are highly contextualized, personalized and hard to grade, such as creative writing, coaching, brainstorming?
* 当我们在现实中部署模型时,无法无限增长测试时思考,我们如何将性能提升平滑地迁移回基础模型,同时降低推理时间成本(例如通过蒸馏)?
* When we deploy the model in reality, we cannot grow test time thinking forever and how can we smoothly translate the performance gain back into the base model with reduced inference time cost (e.g. via distillation)?
* 如何根据手头问题的难度使测试时开销更具适应性?
* How to make test time spending more adaptive according to the difficulty of the problem in hand?