缩放定律,审慎解读

Scaling Laws, Carefully

翁荔 Lilian Weng · Thinking Machines Lab · 2026-06-24 · Lil'Log ↗

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

本文回顾了深度学习中的经验缩放定律,这些定律描述了训练损失如何随模型规模、数据集规模和计算量的增加而按幂律关系可预测地下降。文章追溯了从 Amari 等人的早期理论工作、Hestness 等人的实证研究,到 Kaplan 等人对语言模型有影响力的缩放定律,再到修订计算最优分配的 Chinchilla 论文的历史发展。核心争论围绕如何在模型参数和训练数据之间最优分配计算量:Kaplan 等人认为模型规模的增长应快于数据,而 Chinchilla 通过更仔细的实验和三种互补方法得出结论,模型规模和数据应以相同速率扩展,这意味着许多现有大型模型训练不足。文章强调,缩放定律为预测性能和指导大规模模型训练中的资源分配提供了实用框架,关键要点是模型规模与数据之间的最优平衡对效率至关重要。

This article reviews the empirical scaling laws in deep learning, which describe how training loss decreases predictably with increases in model size, dataset size, and compute, following power-law relationships. It traces the historical development from early theoretical work by Amari et al. and empirical studies by Hestness et al., through the influential Kaplan et al. scaling laws for language models, to the Chinchilla paper that revised the compute-optimal allocation. The core debate centers on how to optimally allocate compute between model parameters and training tokens: Kaplan et al. suggested model size should grow faster than data, while Chinchilla, using more careful experiments and three complementary methods, concluded that model size and data should scale at equal rates, implying that many existing large models were undertrained. The article highlights that scaling laws provide a practical framework for predicting performance and guiding resource allocation in large-scale model training, with the key takeaway being that the optimal balance between model size and data is crucial for efficiency.

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 12)

全文 · Full text(逐段中英对照)

概述 Overview

缩放定律是深度学习中最关键的实证发现之一。其观察形式简单:随着模型规模、数据集规模和算力的扩大,训练损失可预测地下降,遵循幂律曲线,在对数-对数坐标图上呈现为一条直线。我们可以将缩放定律视为描述算力、损失、模型规模和数据之间关系的框架;其核心在于如何在模型规模和数据之间最优地分配宝贵的算力。

Scaling laws are one of the most critical empirical findings in deep learning. The observation is simple in form: the training loss decreases predictably as we scale up model size, dataset size, and compute, following a power-law curve, which appears as a straight line on a log-log plot. We can view scaling laws as a framework for describing the relationship between compute, loss, model size, and data; at its core, it is about how to allocate precious compute optimally between model size and data.

这种可预测性使得缩放定律在实践中极具价值。一种常见的工作流程是:在少量小规模运行上拟合缩放定律,然后外推以估算更大模型的词元(token)和算力需求。

This predictability makes scaling laws highly valuable in practice. A common workflow is to fit scaling laws on a handful of small runs and then extrapolate to estimate the token and compute requirements for larger models.

早期:机器学习损失的可预测性 Early days: ML loss predictability

在缩放定律成为主流概念之前,人们就已经研究了泛化误差随规模变化的可预测性。

The predictability of generalization error with scale had already been investigated before scaling laws became a mainstream concept.

Amari 等人(1992)使用贝叶斯方法和退火近似推导了四种类型的学习曲线。

Amari et al. (1992) derived four types of learning curves using a Bayesian approach and the annealed approximation.

1. 确定性学习算法,无噪声数据,唯一解:\(E = c \cdot N^{-1}\),其中\(c\)为常数。

1. Deterministic learning algorithm, noiseless data, one unique solution: \(E = c \cdot N^{-1}\), where \(c\) is some constant.

2. 确定性学习算法,无噪声数据,多个等价解:\(E = c \cdot N^{-1}\);随着每个新数据点的加入,学习速度更快,因为模型只学习参数的最优流形,而不是寻找单一解点。

2. Deterministic learning algorithm, noiseless data, multiple equivalent solutions: \(E = c \cdot N^{-1}\); the learning is faster with each new data point, because the model only learns the optimal manifold of parameters, instead of finding the single solution point.

3. 确定性学习算法,有噪声数据:\(E = c \cdot N^{-1/2}\);数据中的噪声使学习更加困难。

3. Deterministic learning algorithm, noisy data: \(E = c \cdot N^{-1/2}\); noises in data make learning harder.

4. 随机学习算法,噪声数据:这里的不可约损失是随机学习器无法进一步降低的残差误差,例如当模型在大数据上耗尽容量时。所有四种类型的学习曲线都遵循幂律:

4. Stochastic learning algorithm, noisy data: here the irreducible loss is the residual error that a stochastic learner cannot reduce further, for example when the model runs out of capacity on large data. All four types of learning curves follow a power law:

其中 \(\epsilon\) 可以为 0,\(\beta > 0\)。尽管他们的理论设置基于简化的二分类任务,但为构建经验性的机器学习损失预测模型指明了有用的方向。

where \(\epsilon\) can be 0 and \(\beta > 0\). Although their theoretical setup is based on a simplified binary classification task, it points in a useful direction for building empirical ML loss prediction models.

Hestness 等人(2017)最早的经验研究之一解释了泛化误差、模型大小和数据之间的关系。对于给定的训练数据大小,他们通过网格搜索确定最佳拟合的模型大小,然后绘制损失与训练数据集大小的关系图。在深度学习的四个不同领域(神经机器翻译、图像分类、语言建模和语音识别)中,观察到一种反复出现的模式:

One of the earliest empirical studies by Hestness et al. (2017) explained the relationship between generalization error, model size, and data. For a given training data size, they identified the best-fit model size via grid search and then plotted loss against training dataset size. Across four different domains in deep learning (neural machine translation, image classification, language modeling, and speech recognition), a recurring pattern was observed where:

* 泛化误差随一组因素(如数据大小)呈幂律缩放。

* Generalization error scales as a power law across a set of factors (e.g., data size).

* 模型改进会移动误差曲线,但似乎不影响幂律指数。

* Model improvements shift the error curve but do not seem to affect the power-law exponent.

* 有趣的是,架构改变了幂律拟合的偏移量(\(\\)),但不改变指数(\(\\))。幂律的斜率似乎是问题域的一个属性,而不是模型架构的属性。

* Interestingly, architecture changes the offset (\(\\)) of the power-law fit but does not change the exponent (\(\\)). The slope of the power law appears to be a property of the problem domain rather than the model architecture.

* 拟合大小为的数据集所需的模型参数数量也遵循幂律缩放。

* The number of model parameters needed to fit a dataset of size also scales as a power law.

(左)Deep-Speech-2(DS2)和注意力语音模型的学习曲线,以及(右)不同规模的 DS2 模型的学习曲线。当训练数据变大时,小模型的损失会趋于平稳。(图片来源:Hestness 等人,2017)

Learning curves for (Left) Deep-Speech-2 (DS2) and attention speech model and for (Right) DS2 models of various sizes. The losses of small models plateau when training data becomes large. (Image source: Hestness et al. 2017)

一个概念性图示将学习曲线分为三个阶段。在小数据区域,当学习信号不足时,模型的表现仅略好于随机猜测。在中间(“幂律区域”),我们观察到损失、数据和模型大小之间的幂律关系。最后的不可约误差区域可归因于数据中的噪声等因素。

A conceptual illustration breaks the learning curve into three stages. In the small-data region, when there are not enough learning signals, the model performs only slightly better than random guessing. In the middle (“power-law region”), we observe a power-law relationship between loss, data, and model size. The final irreducible-error region can be attributed to factors such as noise in the data.

幂律学习曲线阶段的图示。(图片来源:Hestness 等人,2017)

Illustration of power-law learning curve phases. (Image source: Hestness et al. 2017)

Rosenfeld 等人(2020)进一步推进了这一方向,尝试将误差建模为模型大小和数据大小的联合函数,并在一系列不同的架构(ResNet、WRN、LSTM、Transformer)和优化器(Adam、SGD 变体)上进行验证。他们通过实验观察到,当固定一个轴时,误差随另一个轴呈幂律衰减:

Rosenfeld et al. (2020) pushed this further by trying to model error as a joint function of both model size and data size, across a diverse set of architectures (ResNet, WRN, LSTM, Transformer) and optimizers (Adam, SGD variants). Empirically they observed that, holding one axis fixed, the error decays as a power law in the other:

其中 \(a, b, c, d\) 是标量常数,\(e\) 不依赖于 \(N\) 或 \(D\)。

where \(a, b, c, d\) are scalar constants and \(e\) is not dependent on either \(N\) or \(D\).

数据大小、模型大小和泛化误差在对数-对数-对数尺度下的三维等高线图。蓝点来自经验实验,曲面是蓝点之间的线性插值。(图片来源:Rosenfeld 等人,2020)

A 3D contour plot of data size, model size and generalization error in log-log-log scale. Blue dots are derived from empirical experiments and the surface is a linear interpolation between blue dots. (Image source: Rosenfeld et al. 2020)

因此,他们可以构建一个简单参数函数形式的预测模型(其中 \(e = 0\)),仅通过在较小的训练配置(\(N, D\) 低于某个阈值)上训练,来预测超过某个阈值时的预期损失。

Thus, they can build a prediction model in the form of a simple parametric function with \(e = 0\) to predict the expected loss for \(N, D > \) certain thresholds by only training on a set of smaller training configs, \(N, D < \) certain thresholds.

在小规模配置上拟合参数化误差模型,并外推到更大的模型/数据规模:(a)实验设置示意图;(b)ImageNet、(c)WikiText-103 和(d)CIFAR100 上的实验结果,使用三种架构(WRN、VGG、DenseNet)和两种优化器(SGD、Adam)进行误差估计。(图片来源:Rosenfeld 等人,2020)

Fitting the parametric error model on small-scale configurations and extrapolating to larger model/data regimes: (a) Illustration of the experiment setup; Experiment results on (b) ImageNet, (c) WikiText-103 and (d) CIFAR100 Error estimation with three architectures (WRN, VGG, DenseNet) and two optimizers (SGD, Adam). (Image source: Rosenfeld et al. 2020)

附注:这些早期工作依赖于经典学习理论的直觉,如 VC 维(模型能够打碎的最大点集的基数)作为容量的代理,但在现代深度学习工作中,VC 维往往过于粗糙,无法解释行为,而经验幂律被证明比理论提供的最坏情况界限更简洁、更实用。

Side note: These early works lean on classical learning-theory intuition like the VC dimension (the cardinality of the largest set of points a model can shatter) as a proxy for capacity, but in modern deep learning work the VC dimension is often too coarse to explain the behavior and the empirical power laws turned out to be much cleaner and more practical than the worst-case bounds that theory provides.

Kaplan 等人的缩放定律 Kaplan et al.’s Scaling Laws

Kaplan 等人(2020)在语言建模社区推广了缩放定律的概念。他们发现交叉熵测试损失分别与模型大小(不包括嵌入层)、数据集大小和训练算力在多个数量级上呈幂律关系。这些发现与上一节中的早期工作一致,但 Kaplan 等人通过关注 Transformer 语言模型和更大规模的实证实验,将这一概念形式化,其中模型大小从 768M 到 1.5B 非嵌入参数,数据集大小从 22M 到 23B 个词元。论文中的所有训练运行都使用了学习率调度,包括 3000 步的线性预热,随后是余弦衰减至零。

Kaplan et al. (2020) popularized the concept of scaling laws in the language modeling community. They found that the cross-entropy test loss scales as a power law with each of model size (excluding embedding layers), dataset size, and training compute across many orders of magnitude. The findings are aligned with early work in the last section, but Kaplan et al. formalized the concept with a focus on Transformer language models and empirical experimentation at a larger scale, with model size ranging from 768M to 1.5B non-embedding parameters and dataset size from 22M to 23B tokens. All training runs in the paper used a learning rate schedule with a 3000 step linear warmup, followed by a cosine decay to zero.

* 损失分别与模型大小、数据集大小和算力呈幂律关系;为了获得最佳性能,三者必须同步扩展。

* The loss scales as a power law with model size, dataset size, and compute individually; for optimal performance all three must scale in tandem.

* 训练曲线遵循可预测的幂律,其参数大致与模型大小无关。

* Training curves follow predictable power laws whose parameters are roughly independent of model size.

* 更大的模型样本效率更高,这意味着它们比小模型用更少的优化步骤和更少的数据点就能达到给定的损失。

* Larger models are more sample-efficient, meaning that they reach a given loss with fewer optimization steps and fewer data points than small models.

* 架构细节(宽度、纵横比等)不如纯粹的规模重要。

* Architectural details (width, aspect ratio, etc.) matter less than sheer scale.

* 训练损失与测试损失正相关。(听起来微不足道,但这是预训练工作的基础。另一方面,预训练损失的改善是否能迁移到后训练评估中,还需要单独研究。)

* Train loss and test loss are positively correlated. (Sounds trivial but this is the foundation for pretraining work. On the other hand, whether pretraining loss improvement transfers to posttraining evaluation needs separate studies.)

* 在给定算力预算的情况下,训练一个非常大的模型并在 _收敛之前_ 停止,比训练一个较小的模型直到收敛更高效。这一发现正是 Chinchilla 缩放定律(下一节)所不同意的:Kaplan 等人高估了最优模型大小,因为他们拟合的指数偏大。

* Given a fixed compute budget, it is more efficient to train a very large model and stop _before convergence_ than to train a smaller model all the way to convergence. This finding is where the Chinchilla scaling laws (the next section) disagree: Kaplan et al. overestimated the optimal model size as their fitted exponent was larger.

他们将对 和 的联合依赖关系总结为一个方程:

They summarize the joint dependence on and in a single equation:

这种形式的一个良好结果是,过拟合的程度(即模型复杂或数据量小)主要取决于比值 ,这表明数据需要以与模型大小增长成特定比例的方式增长,以避免训练受数据限制。

A nice consequence of this form is that the extent of overfitting (i.e. model is complex or data is small) depends predominantly on the ratio , which indicates that the data needs to grow in a specific proportion to the growth of the model size to avoid training being data-limited.

测试损失作为算力、数据集大小和参数的幂律,跨越多个数量级。(图片来源:Kaplan 等人,2020)

Test loss as a power law in compute, dataset size, and parameters, spanning many orders of magnitude. (Image source: Kaplan et al. 2020)

最具影响力、事后看来也最具争议的结论是算力最优分配。Kaplan 等人发现并得出结论:模型规模的增长速度应快于数据集规模。具体而言,对于 10 倍的算力增加,他们建议将模型规模扩大约 5.5 倍,而训练 token 数仅增加约 1.8 倍。后来的 Chinchilla 论文推翻了这一建议,认为该建议会导致大模型严重“欠训练”。

The most influential and, in hindsight, most contested conclusion was the compute-optimal allocation. Kaplan et al. found and concluded that model size should grow faster than dataset size. Concretely, for a 10x increase in compute they suggested scaling the model size by ~5.5x but the training tokens by only ~1.8x. The Chinchilla paper would later overturn this recommendation, arguing that it leaves large models badly _undertrained_.

Kaplan 等人的另一项有用分析基于 和 近似估计了所需的训练 FLOPs 数量。每次乘加运算计为约 2 FLOPs。

Another useful analysis in Kaplan et al. approximates the number of training FLOPs needed based on and . Each multiply-add is counted as ~2 FLOPs.

针对不同 Transformer 架构组件的参数和算力估计,给定层数 、模型宽度(= ;原表中符号不一致)、前馈层维度(通常等价于 )、注意力维度(通常等价于 )、上下文长度 和词汇表大小 。(图片来源:Kaplan 等人,2020)

Parameter and compute estimation for different Transformer architectural components, given the number of layers , model width (= ; the notation is inconsistent in the original table), dimension of feed-forward layer (often equivalent to , attention dimension (often equivalent to ), the context length and the vocabulary size . (Image source: Kaplan et al. 2020)

给定标准配置,其中 ,并且从 和每 token 前向计算中排除嵌入层:

Given a standard config where , and excluding embedding layers from and the per-token forward compute:

然后我们将反向传播的 FLOPs 计为前向传播 FLOPs 的两倍,因为反向传播执行两次矩阵乘法,分别计算关于输入激活和权重的梯度。因此,每个 token 的训练 FLOPs 总计约为 ,在 个 token 上训练的总 FLOPs 为 。

Then we count backward-pass FLOPs as twice the forward-pass FLOPs, because backpropagation runs two matrix multiplications, for gradients with respect to the input activations and the weights, respectively. Thus, in total, the training FLOPs per token are approximately , and the total FLOPs for training over tokens are .

Chinchilla 缩放定律 Chinchilla Scaling Laws

Chinchilla 论文(Hoffmann 等人,2022)通过更精细的实验设计,研究了在固定算力预算下,最优模型规模(总参数,_包括_嵌入层)与训练 token 数量之间的关系,得出了与 Kaplan 等人略有不同的结论。

The Chinchilla paper (Hoffmann et al. 2022) studied the relationship between the optimal model size (total parameters, _including_ embeddings) and the number of tokens under a _fixed_ compute budget with a more careful experimental design and arrived at a somewhat different answer from Kaplan et al..

你应该知道 Chinchilla 长什么样 😊(图片来源:ChatGPT 生成)

You should know how chinchilla looks 😊 (Image source: ChatGPT generated)

核心问题是在给定约束下如何最优地分配资源。换句话说,当我们只有有限的 FLOPs(即一定数量的 GPU 运行一定时间)时,我们应该如何在更多的数据 token 和更大的模型参数之间进行选择?

The central question is on the best strategy to allocate resources given a constraint . In other words, when we have only limited FLOPs (a given number of GPUs running for a given period of time), how should we choose between more data tokens and more model parameters?

Chinchilla 论文提出了三种设计精巧的缩放定律拟合方法。

The Chinchilla paper presented three neatly designed methods for scaling laws fitting.

实证实验扫描了 400 多个模型,规模从 70M 到超过 16B 参数,训练 token 从 5B 到 500B。实验假设每个训练 token 都是唯一的(无限数据场景)。所有运行都采用余弦学习率调度,在训练周期内衰减 10 倍。通过扫描模型规模,描绘出计算最优前沿。

The empirical experiments scanned over 400 models, with sizes from 70M to over 16B parameters and training tokens from 5B to 500B. The experiments were under the assumption that every training token is unique (the infinite-data regime). All runs used a cosine learning-rate schedule decaying by 10x over the training horizon. Sweeping over model sizes traces out the compute-optimal frontier.

方法 1:固定模型规模,变化词元预算 Method 1: Fix model sizes, vary the token budget

对于每个参数数量,使用不同的词元预算进行多次训练,并记录每个 FLOP 预算下达到的最小损失。

For each parameter count, train several runs with different token budgets, and record the minimal loss achieved per FLOP budget.

Chinchilla 方法 1:不同模型规模下,训练损失随 FLOP 预算变化的曲线。(图片来源:Hoffmann 等人,2022)

Chinchilla Method 1: training loss curves over FLOP budgets for a sweep of model sizes. (Image source: Hoffmann et al. 2022)

方法 2:IsoFLOP 曲线 Method 2: IsoFLOP profiles

固定算力预算,绘制最终损失随参数数量的变化曲线。每条等 FLOP 曲线在对数空间中近似为抛物线,其最小值标志着该算力预算下的最优模型规模。随后,对不同预算重复此过程,在图中可得到一条幂律线。

Fix a compute budget and plot the final loss against parameter count. Each iso-FLOP curve is roughly a parabola in log-space, and its minimum flags the optimal model size for that compute budget. Then repeating across budgets traces a power-law line in the plot.

Chinchilla 方法 2:等 FLOP 抛物线;每条曲线的最小值即为该预算下的算力最优模型规模。(图片来源:Hoffmann 等人,2022)

Chinchilla Method 2: IsoFLOP parabolas; the minimum of each curve is the compute-optimal model size for that budget. (Image source: Hoffmann et al. 2022)

方法 3:参数拟合 Method 3: Parametric fit

我们实际上可以通过在约束下最小化来获得最优的闭式近似。

We can actually get a closed form approximation of the optimal by minimizing under the constraint .

首先将表达式化简为仅包含 的形式:

First let’s reduce the expression to contain only :

当 时,模型大小和训练 token 应以相同的速率扩展。

When , model size and training tokens should scale at equal rates.

为了找到最优的 ,Chinchilla 论文采用了 Huber 损失(对异常值具有鲁棒性; )和 L-BFGS 算法(适用于参数数量较少的曲线拟合)。

To find the optimal , the Chinchilla paper adopts a Huber loss (robust to outliers; ) and the L-BFGS algorithm (good for curve fitting with a small number of parameters).

Chinchilla 通过三种互补的方法得出了其答案,这些方法的最终结果相互一致,这也是其结果相当有说服力的部分原因。

Chinchilla arrives at its answer through three complementary methods whose final results agree with each other, and this is part of why the result was quite convincing.

这三种方法在计算最优前沿上达成一致,其中 ,但与 Kaplan 等人的结果不一致。注意,方法 3 的结果与其他两种略有偏差,我们将在后面解释。(图片来源:Hoffmann 等人,2022)

The three methods agree on a compute-optimal frontier where , but disagree with Kaplan et al. Note that method 3's results are slightly off from the other two, which we will explain later. (Image source: Hoffmann et al. 2022)

该图展示了三种不同方法的 Chinchilla 预测结果,以及 Kaplan 等人(2020)的预测。所有三种方法都表明,当时的主流大语言模型(LLM)训练不足。(图片来源:Hoffmann 等人,2022)

The plot of the Chinchilla predictions by three different approaches, as well as predictions by Kaplan et al. (2020). All three methods suggest that several mainstream LLMs at the time were undertrained. (Image source: Hoffmann et al. 2022)

Chinchilla 论文中关于大多数大型模型(当时,约 2022 年)训练不足的主张,得到了一个著名演示的支持:在与 Gopher(Rae 等人,2021;280B 参数,300B token 预算)相同的算力预算下,他们训练了 Chinchilla(70B 参数,1.4T token 预算),该模型规模小 4 倍,但训练 token 数约为其 4 倍,并且在各方面都优于 Gopher。

The claim in the Chinchilla paper that most large models (at the time, ~2022) were undertrained is supported by a famous demonstration: under the same compute budget as Gopher (Rae et al. 2021; 280B parameter count, 300B token budget), they trained Chinchilla (70B parameter count, 1.4T token budget), a model 4x smaller but trained on roughly 4x more tokens and it outperformed Gopher across the board.

调和 Kaplan 与 Chinchilla Reconciling Kaplan and Chinchilla

Chinchilla 缩放定律与 Kaplan 等人的结论存在以下分歧:

The Chinchilla scaling laws disagree with Kaplan et al. as follows:

* 不再是“模型规模的增长应快于数据”(\(N \propto D^{0.74}\)),而是模型规模每翻一倍,训练 token 数也应翻倍(\(N \propto D\))。

* Instead of “grow the model faster than the data” (\(N \propto D^{0.74}\)), for every doubling of model size, you should also double the number of training tokens (\(N \propto D\)).

* 不再是“训练一个大模型并在收敛前停止”,而应在更多数据上训练一个较小的模型。

* Instead of “train a big model and stop before convergence,” you should train a smaller model on more data.

两篇论文在基本原则上是相同的,但它们在最优规模与 token 数的权衡点上存在分歧。为什么分歧如此之大?

Both papers still agree on the same underlying principle, but they disagree on where the optimal size-vs-token tradeoff lies. Why do they disagree so much?

差异 1:Kaplan 等人的实验主要针对小模型。Kaplan 等人的实验主要针对较小的模型,而 Chinchilla 论文的实验规模达到了 10 倍以上。在对数-对数空间中进行外推时,拟合上的微小差异可能导致巨大差异(参见玩具模拟)。

Difference 1: Kaplan et al. experimented mostly on small models. Kaplan et al. experimented mostly on smaller models, while the Chinchilla paper’s experiments reached more than 10x larger scales. When we extrapolate in log-log space, a small difference in the fit can result in large differences (See toy simulation).

差异 2:嵌入参数数量对小模型很重要。在小参数范围内,嵌入参数占总参数的比例不可忽略,因此是否计入它们会产生影响。Pearce & Song (2024) 沿着这一思路进行了深入分析。我们用 \(N\) 表示排除嵌入后的模型大小和计算量,用 \(N_{total}\) 表示总参数数量。

Difference 2: Embedding parameter count matters for small models. In the small-parameter regime, embedding parameters are a non-negligible fraction of the total and thus counting them or not matters. Pearce & Song (2024) did a thorough analysis along this line. Let’s use \(N\) to denote model size and compute when embedding is excluded and use \(N_{total}\) to count total parameters.

为了弥合两者,他们拟合了总参数与非嵌入参数之间的关系 \(N_{total} = c N^{\alpha}\),其中 \(c\) 为常数:

To bridge them, they fit a relationship between total parameters and non-embedding parameters \(N_{total} = c N^{\alpha}\), for some constant \(c\):

这种形式具有良好的性质:严格递增,且 \(N_{total} > N\)(因为 \(c > 1\) 且 \(\alpha > 0\))。

This form has nice properties of being strictly increasing and \(N_{total} > N\) (because \(c > 1\) and \(\alpha > 0\).

将其代入 Chinchilla 定律方程,

Plugging this into the Chinchilla laws equation,

上述方程中 \(N\) 与 \(D\) 的关系不再是简单的幂律。我们只能在局部将其近似为 \(D \propto N^{\beta}\),其中 \(\beta\) 是基于一阶导数的局部指数(\(\beta = \frac{d \log D}{d \log N}\)),而非全局幂律指数,因此得到 \(\beta = \frac{1}{\alpha \gamma}\)。关于指数近似的完整细节,请参见 Pearce & Song (2024) 的附录 A.1。

The relationship between \(N\) and \(D\) in the above equation is no longer a clean power law. We can only approximate it locally as \(D \propto N^{\beta}\), where \(\beta\) is a local exponent based on a first-order derivative (\(\beta = \frac{d \log D}{d \log N}\)) rather than a global power-law exponent, resulting in \(\beta = \frac{1}{\alpha \gamma}\). See the full details of how the exponent is approximated in Appendix A.1 in Pearce & Song (2024).

局部幂律指数随 增长的示意图。(图片来源:Pearce & Song 2024)

Visualization of how the local power-law exponent grows with . (Image source: Pearce & Song 2024)

如上图所示,随着 增大, 收敛到 Chinchilla 估计值。通过使用上述方程生成合成训练曲线,在模型规模从 768M 到 1.5B 的范围内(如 Kaplan 等人所述),他们估计该区域内的 接近 Kaplan 系数 0.73。

As shown in the visualization above, as gets larger, converges to the Chinchilla estimate. By generating synthetic training curves using the above equation, in the range of model size from 768M to 1.5B (as in Kaplan et al.), they estimated that is close to the Kaplan coefficient of 0.73 in that region.

为什么是幂律? Why power law?

幂律在 AI 之外的许多领域中被广泛观察到,例如齐普夫定律、无标度网络、城市缩放定律以及许多其他复杂系统。反复出现的模式是:大事件罕见,小事件常见,且大小与频率之间的关系在双对数坐标下通常呈直线。

Power laws are widely observed across many domains outside AI, such as in Zipf's law, scale-free networks, urban scaling laws, and many other complex systems. The recurring pattern is that large events are rare, small events are common, and the relationship between size and frequency often follows a straight line at log-log scale.

为什么 LLM 的缩放定律也具有幂律的形状?

Why do LLM scaling laws also have the shape of a power law?

部分受到不同领域显示出不同指数(Hestness 等人,2017)的启发,Sharma 和 Kaplan(2020)提出的一个早期解释假设,语言建模可以被视为在数据的低维流形上进行回归。更多的模型参数可以导致对数据流形的更精细划分,从而减小泛化误差。简单来说,如果一个有效大小为\(N\)的模型将\(d\)维流形划分为\(N\)个区域,则典型的线性分辨率按\(N^{-1/d}\)缩放。这与上述缩放定律具有类似的幂律形式。该理论在无限数据、欠拟合情况下最为适用,但实际上估计数据流形的内在维度相当困难。

Inspired partly by different domains displaying different exponents (Hestness et al. 2017), one early explanation by Sharma & Kaplan (2020) hypothesizes that language modeling can be viewed as doing regression on a low-dimensional manifold of data. More model parameters can induce a finer partition of the data manifold and therefore smaller generalization error. In the simplest terms, if a model of effective size partitions a \(d\)-dimensional manifold into \(N\) regions, the typical linear resolution scales like \(N^{-1/d}\). This has a similar power-law form to the scaling laws above. This theory applies most cleanly in the infinite-data, underfitting regime, but in reality estimating the intrinsic dimension of a data manifold is quite hard.

后来的一个假设(Michaud 等人,2023;Brill,2024)认为,知识或技能是以离散块(“量化”)的方式学习的,并且这些技能的频率分布遵循幂律。模型先学习常见技能,后学习罕见技能,导致损失呈现平滑的幂律衰减。

A later hypothesis (Michaud et al. 2023, Brill 2024) assumes that knowledge or skills are learned in discrete chunks (“quantized”) and that the frequency distribution of these skills follows a power law. The model learns common skills first and rare skills later, resulting in a smooth power-law decay in loss.

这里我只列出了两个假设,但还有更多研究通过数据的谱尾、核特征值、自然语言统计或训练动态中的相变来解释幂律缩放的形状。

I only listed two hypotheses here, but there are more studies on explaining the shape of power-law scaling through spectral tails of data, kernel eigenvalues, natural-language statistics, or phase transitions in training dynamics.

数据受限区域的缩放定律 Scaling Laws in Data-Limited Region

经典的缩放定律假设有效无限多的唯一数据、无重复且无多轮训练。随着模型规模显著增长,我们正在耗尽足够高质量的唯一 token。事实上,关于 AI 缩放能持续多久的一些争论,核心就在于我们是否正在撞上“数据墙”。

Classic scaling laws assume effectively unlimited unique data, no repetition, and no multi-epoch training. As the model size grows significantly, we are running out of enough high-quality unique tokens. In fact, some arguments about how long scaling in AI can continue are centered on whether we are hitting a “data wall”.

同样值得强调的是,背后的数据集预期已经过清洗。预训练数据流水线往往是高效预训练流程的重要部分,常见步骤包括去重(精确和模糊)、质量过滤、样板移除、安全过滤、PII/版权掩蔽、基准去污染,以及基于语言、质量、内容类型等对数据混合成分的仔细重新加权。即使两个数据集的 token 数相同,高质量数据集与互联网垃圾数据集也可能产生截然不同的算力效率。

It is also worth emphasizing that the dataset behind is expected to be already cleaned. The pretraining data pipeline is often a large part of an effective pretraining pipeline, with common steps like deduplication (exact and fuzzy), quality filtering, boilerplate removal, safety filtering, PII/copyright masking, benchmark decontamination and careful reweighting of data mix components based on language, quality, content type, etc. Even when two datasets contain the same token count, a high-quality dataset and a dataset of Internet slop can yield drastically different compute efficiency.

Hernandez 等人(2022)的研究聚焦于一个受控版本:一个基本唯一的数据集,其中包含少量重复数据。从大型数据集出发,数据混合保留 90% 的非重复数据,但将剩余 10% 替换为原始数据中极小部分的重复。通过训练一个 Transformer 模型 100B token,他们观察到双重下降现象,即测试损失实际上会随着重复数据强调程度的增加而先变差再变好,且该效应随着重复比例的增加而更加显著。

The study by Hernandez et al. (2022) focused on a controlled version: a mostly-unique dataset with a small fraction of repeated data. Starting from a large dataset, the data mix keeps 90% non-repeated but replaces the remaining 10% with repeats of a tiny portion of the original. By training a Transformer model for 100B tokens, they observed a double-descent phenomenon, that is, the test loss can actually get worse and then better again as a function of how much the repeated data is emphasized, an effect that becomes more pronounced as the repeated fraction grows.

测试损失随重复比例增加而出现的双重下降(左侧重复 90%,右侧 50%)。(图片来源:Hernandez 等人,2022)

Double-descent in the test loss as the repeated fraction increases (90% repeated on the left, 50% on the right). (Image source: Hernandez et al. 2022)

训练中期的平坦或上升趋势可能归因于对重复数据的记忆。具有此类形状的学习曲线使得缩放定律拟合不够准确。他们还得出结论,重复数据会损害某些 OOD 评估和下游微调。然而,他们的数据混合是在更类似实验室的设置中构建的,而真实世界数据中的重复往往更加微妙(例如,不同数据具有不同级别的重复、语义重复等)。

The flat or increasing trend in the middle of training is possibly due to memorization of repeated data. Learning curves with such shapes make scaling law fitting less accurate. They also concluded repeated data hurts some OOD evaluation and downstream fine-tuning. However, their data mix is constructed in a more lab-like setup, and repetition in real-world data is often more nuanced (e.g. different data has different levels of repetition, semantic repetition, etc.).

与其说数据重复会损害训练,我们更感兴趣的是如何在高质量独特数据并非无限、且训练中可能需要重复数据的情况下拟合缩放定律。

Rather than saying data repetition hurts training, we are more interested in how to fit scaling laws, given that the unique high-quality data is not infinite and we likely have to repeat data during training.

Muennighoff 等人(2023)研究了在数据受限条件下如何最优分配算力的研究问题。具体而言,他们通过约 400 次实验,参数规模从 10M 到 9B,数据规模高达 900B tokens,训练轮数最多 1500 轮,实证研究了数据重复的影响。每个 epoch 重复完全相同的数据集,并在 epoch 之间进行洗牌,在留出测试集上进行评估。

Muennighoff et al. (2023) took on the research question of how compute should be allocated optimally when model training is data-constrained. Specifically, they empirically studied the impact of data repetition across roughly 400 experiments, 10M–9B parameters, data sizes up to 900B tokens, and up to 1500 epochs. The exact same dataset is repeated each epoch, shuffled between epochs, and evaluated on a held-out test set.

关键的建模调整是将总 token 数分解为两部分:(i)独特 token 的数量和(ii)重复次数(即 epoch 数减 1)。因此我们有\(D = U + R\)。在独特数据预算\(U\)下,根据定义\(R = D - U\)和\(D = U + R\)。他们使用 Chinchilla 缩放定律来找到拟合\(D\)的最优模型大小,并通过重复次数\(R\)定义多余模型大小。

The key modeling adjustment is to decompose the total token count into two parts: (i) the number of unique tokens and (ii) the number of repeats (i.e. num. epochs - 1). Thus we have \(D = U + R\). With a unique-data budget \(U\), by definition \(R = D - U\) and \(D = U + R\). They use the Chinchilla scaling laws to find the optimal model size for fitting \(D\), and define excess model size via repeats \(R\).

然后,他们更新了 Chinchilla 参数拟合(方法 3),使用有效(折扣后的)数据和模型大小代替原始量:

They then update the Chinchilla parametric fit (method 3) to use effective (discounted) data and model size in place of the raw quantities:

直觉是,token 的价值随着重复而呈指数衰减。在他们的建模中,每次重复都会消耗 token 剩余价值的一部分,其中\(λ\)是可学习的“半衰期”参数。当\(λ = 0\)或\(λ = 1\)时,我们恢复\(D_{\text{eff}} = D\)。

The intuition is that a token’s value decays _exponentially_ as it is repeated. In their modeling, each repetition costs the token a fraction of its remaining value, where \(λ\) is a learnable “half-life” parameter. When \(λ = 0\) or \(λ = 1\), we recover \(D_{\text{eff}} = D\).

一种对称的公式处理了模型规模过剩的情况,体现了“更大的模型在重复数据上过拟合更快”以及“模型可能相对于其数据集过大”的思想。这一部分不太直观,我未能找到令人满意的解释来说明为什么模型规模需要以与重复数据对称的形式出现。后来 Lovelace 等人(2026)的工作改变了这一假设。

A symmetric formulation handles excess model size, capturing the idea that “larger models overfit more quickly on repeated data” and that “a model can be too large for its dataset.” This component is less intuitive, and I could not find a satisfactory explanation for why model size needs to appear in such a symmetric form as repeated data. Later work by Lovelace et al. (2026) changed this assumption.

他们的经验拟合发现,_过剩参数在数值上比重复数据衰减得更快_,因此我们应该将更多资源分配给更多轮次,而不是更多模型参数。正如作者所指出的,这种建模的一个弱点是它显著低估了失败模型(即训练中途损失上升的模型)的最终测试损失,例如训练了 44 轮的模型。

Their empirical fit finds that _excess parameters decay faster in value than repeated data_, so we should allocate more resources on more epochs rather than more model parameters. One weakness of this modeling, as the authors also pointed out, is that it significantly underestimates the final test loss of failing models (i.e. models whose loss increases midway through training), such as models trained for 44 epochs.

在重复条件下的数据受限缩放比不考虑数据的拟合更好地捕捉了实验结果;重复词元的价值呈指数衰减并趋于上限。随着轮次增加,拟合效果变差,因为高重复导致测试损失在训练中途上升,这在图中未显示。(图片来源:Muennighoff 等人,2023)

Data-constrained scaling under repetition captures the experimental results better than data-unaware fitting; the value of repeated tokens decays exponentially toward a ceiling. The fitting gets worse with more epochs as high repetition causes the test loss to increase midway through training, not depicted in the plot. (Image source: Muennighoff et al. 2023)

最近,Lovelace 等人(2026)用不同的方法重新审视了同一问题。他们没有将过参数化建模为有效模型规模的收益递减,而是显式地建模了模型规模与数据重复之间的交互。经验上,他们训练了约 300 个模型,参数规模从 1500 万到 10 亿,独特词元从 5000 万到 60 亿。

Most recently, Lovelace et al. (2026) revisited the same problem with a different approach. Rather than modeling overparameterization as a diminishing return on effective model size, Lovelace et al. model the interaction between model size and data repetition explicitly. Empirically, they trained about 300 models, spanning 15M to 1B parameters and 50M to 6B unique tokens.

当他们绘制固定模型规模在不同数据重复水平下的拟合残差时,观察结果是直观的:更多轮次造成更大损害,有趣的是_更大的模型对重复更敏感_。这暗示损失惩罚可能是模型规模和数据集规模的函数。

When they plot the fit residual for a fixed model size across a range of data-repetition levels, the observation is intuitive: more epochs cause more damage, and interestingly _larger models are more sensitive_ to repetition. This hints that the loss penalty is likely a function of both model size and data size.

有效规模拟合的残差表明,过拟合损害随训练轮数和模型规模的增加而加剧。(图片来源:Lovelace 等人,2026)

Residuals of the effective-size fit reveal that overfitting damage grows with both the number of epochs and the model size. (Image source: Lovelace et al. 2026)

我们引入了一个显式的过拟合惩罚项,该惩罚项围绕**容量比**(参数数量相对于唯一 token 的数量)构建:

An explicit overfitting penalty term was introduced and built around the _capacity ratio_ (parameter count relative to unique tokens):

* 指数(第二个可学习参数)使惩罚项随容量比非线性缩放;

* the exponent (the 2nd learnable parameter) lets the penalty scale nonlinearly with the capacity ratio ;

* 重复次数上的独立指数(第三个可学习参数)将重复非线性与容量比解耦。

* the separate exponent (the 3rd learnable parameter) on the repetition count decouples repetition nonlinearity from .

新增项(红色部分)是一个直接的过拟合惩罚,它随数据重复次数和模型相对于可用唯一数据的过参数化程度的增加而增长。

The added term (in red) is a direct overfitting penalty that grows with both how many times you repeat the data and how over-parameterized the model is relative to the unique data available.

他们还进行了一项关于权重衰减在有限数据约束下如何影响训练的案例研究,发现强权重衰减能减少数据重复引起的过拟合惩罚。

They also conducted a case study on how weight decay affects training under the limited-data constraint, finding that strong weight decay reduces the overfitting penalty caused by data repetition.

强权重衰减能减少数据重复引起的过拟合惩罚。(图片来源:Lovelace 等人,2026)

Strong weight decay reduces the overfitting penalty from data repetition. (Image source: Lovelace et al. 2026)

Muennighoff 等人和 Lovelace 等人的建模方法都基于经验曲线拟合,因此尚不清楚为什么数据受限的缩放定律应具有这些确切形式,以及为什么需要每个自由参数。我们期待更多沿此方向的理论工作。

Both modeling approaches by Muennighoff et al. and Lovelace et al. are based on empirical curve fitting, so it remains unclear why data-constrained scaling laws should take exactly these forms and why each free parameter is necessary. We are curious about more theoretical work along this line.

现实中拟合缩放定律的棘手之处 Trickiness of Fitting Scaling Laws in Reality

尽管其形式简洁,但在实践中,缩放定律的拟合可能对看似微不足道的程序性选择异常敏感,例如如何统计参数、如何舍入精度、如何对损失求和或取平均等。

Despite its clean form, in practice, scaling law fitting can be surprisingly sensitive to seemingly trivial procedural choices, like how you count parameters, how you round the precision, how you sum or average the loss, etc.

由于缩放定律仅适用于我们能够负担训练的相对较小、相对便宜的模型,而预测是针对大数个数量级的模型进行外推。在这样的设置下,看似舍入误差的选择可能导致预测结果出现巨大差异。

Because a scaling law is only fit on the (relatively small, relatively cheap) models that we can afford to train, and the prediction is _extrapolated_ for a model orders of magnitude larger. In such a setup, choices that look like rounding error may lead to wild differences in prediction.

同时,缩放定律拟合假设唯一变化的因素是规模,这意味着模型架构、优化器、学习率调度、批量大小提升、数据混合、分词器及其他设计选择应保持不变。另一个潜在假设是所有这些设置都应经过仔细调优,因为欠训练模型等情况可能导致不同的结论。

Meanwhile, scaling-law fitting assumes the only changing factor is _scale_, which means that the model architecture, optimizer, learning rate schedule, batch ramp, data mix, tokenizer, and other design choices should remain the same. Another underlying assumption is that all these settings should have been carefully tuned, as cases like undertrained models can lead to a different conclusion.

Kaplan 等人与 Chinchilla 的结果不一致是展示缩放定律拟合棘手性的一个例子。

The disagreement between results by Kaplan et al. and Chinchilla is one example to showcase the trickiness of scaling laws fitting.

* L-BFGS-B 最小化器中的高损失尺度,是由于对示例的 Huber 损失值取平均而非求和所致,这导致优化提前终止。在原始拟合和自举过程中,损失最小化的提前停止产生了不一致的估计和难以置信的窄置信区间。

* A high loss scale in the L-BFGS-B minimizer, caused by averaging Huber-loss values over examples instead of summing them, which led to premature termination of the optimization. The early stopping of loss minimization during both the original fit and bootstrapping produced inconsistent estimates and implausibly narrow confidence intervals.

* 所报告的参数值被四舍五入到 2 位有效数字,这使得推导出的缩放定律看起来比实际更不准确。

* The reported values of the parameters were rounded to 2 digits of precision, which made the derived scaling laws appear less accurate than they actually were.

玩具模拟 Toy simulation

这里是一个由 ChatGPT 创建的玩具模拟小部件,旨在演示三种特定的失败模式。

Here is a toy simulation widget, created by ChatGPT, designed to demonstrate three specific failure modes.

因此 。这是 Besiroglu 等人(2024)的估计。

and thus . This is the estimate from Besiroglu et al. (2024).

该模拟绘制了损失预测与数据集大小的关系图,并提供了一组滑块来展示:

The simulation plots the loss prediction vs dataset size , while providing a set of sliders to show case:

* 损失精度:将损失从高精度舍入到低精度小数位会改变拟合参数值。

* Loss precision: rounding losses from high to low decimal points can change the fitted parameter values.

* 损失噪声:仅以毫损失(0.001)单位的倍数扰动损失值就会导致不同的拟合结果。

* Loss noise: perturbing loss values by only a multiplier of milli-loss (0.001) units leads to different fit.

* 拟合区域敏感性:仅拟合小模型、仅拟合中等模型或拟合所有模型,会得到不同的表观缩放定律。

* Fit-region sensitivity: fitting only small models, only medium models, or all models gives different apparent scaling laws.

互动版:图/公式 + 针对本篇提问 →