规模假设

The Scaling Hypothesis

Gwern Branwen Gwern Branwen · · 2020-05-28 · Gwern.net ↗

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

跳转到主要内容。关于 GPT-3:元学习、规模、影响和深层理论。规模假设:神经网络吸收数据和计算,随着问题变难而泛化并变得更贝叶斯,即使在按全球标准微不足道的规模下也能展现新能力。深度学习革命正如预言般开始。

Skip to main content[](https://gwern.net/index) GPT-3, AI scaling, algorithm, insight porn, AI safety, RL scaling, sociology, transhumanism On GPT-3: meta-learning, scaling, implications, and deep theory. The scaling hypothesis: neural nets absorb data & compute, generalizing and becoming more Bayesian as problems get harder, manifesting new abilities even at trivial-by-global-standards-scale. The deep learning revolution has begun as foretold.

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 20)

全文 · Full text(逐段中英对照)

概述 Overview

跳转到主要内容[](https://gwern.net/index)

Skip to main content[](https://gwern.net/index)

GPT-3、AI Scaling(规模扩张)、算法、洞察色情、AI 安全、RL Scaling(强化学习规模扩张)、社会学、超人类主义

GPT-3, AI scaling, algorithm, insight porn, AI safety, RL scaling, sociology, transhumanism

论 GPT-3:元学习、Scaling(规模扩张)、影响及深层理论。Scaling 假设(缩放假设):神经网络吸收数据与算力,随着问题变难而泛化并变得更贝叶斯,即使在按全球标准微不足道的规模下也能展现新能力。深度学习革命正如预言般开始。

On GPT-3: meta-learning, scaling, implications, and deep theory. The scaling hypothesis: neural nets absorb data & compute, generalizing and becoming more Bayesian as problems get harder, manifesting new abilities even at trivial-by-global-standards-scale. The deep learning revolution has begun as foretold.

2020-05-28–2022-01-02 _完成_⁠确定性:_可能_⁠重要性:_10_⁠反向链接⁠⁠类似⁠⁠参考文献⁠

2020-05-28–2022-01-02 _finished_⁠certainty: _likely_⁠importance: _10_⁠backlinks⁠⁠similar⁠⁠bibliography⁠

论“GPT-3:语言模型是少样本学习者”,Brown 等人 2020⁠(诗歌⁠及我的后续“GPT-3 创意写作”⁠,比较我旧的微调 GPT-2 诗歌⁠;⁠随机样本;“OpenAI API”⁠及真实世界演示)

On “GPT-3: Language Models are Few-Shot Learners”, Brown et al 2020⁠ (poems⁠& my followup “GPT-3 Creative Writing”⁠, compare my old finetuned GPT-2 poetry⁠; ⁠random samples; “OpenAI API”⁠ with real-world demos)

我强烈建议任何对 GPT-3 感兴趣的人也至少浏览一下 OA 的⁠随机样本,或者更好的是,阅读我在“GPT-3 创意写作”中的样本——阅读论文和查看一些标准基准图表并不能很好地感受与 GPT-3 合作的感觉,也无法体会它所能做的各种事情,而这些是基准测试所遗漏的。

I strongly encourage anyone interested in GPT-3 to also at least skim OA’s ⁠random samples, or better yet, my samples in “GPT-3 Creative Writing”—reading the paper & looking at some standard benchmark graphs does not give a good feel for what working with GPT-3 is like or the diversity of things it can do which are missed by benchmarks.

元学习 Meta-Learning(https://gwern.net/scaling-hypothesis#meta-learning "Link to section: § 'Meta-Learning'")

学会学习。2020 年 5 月,OpenAI 发布了 GPT-2 的长期后续模型——一个统治一切的模型,但令人惊讶的是,研究人员对此兴趣寥寥,没有博客文章,没有媒体宣传,除了轻蔑的否定之外,几乎没有公开讨论。该模型比 GPT-2 大 117 倍,拥有 1750 亿参数,具备更强大的语言生成能力,使其能够解决从算术、英语翻译、解谜到 SAT 类比推理等各种问题——完全通过文本示例进行提示,无需任何专门训练或微调,仅仅基于在大型互联网文本语料库上的下一个词预测训练。这意味着 GPT-3 的注意力机制充当了“快速权重”,通过在足够多样的数据上训练而“学会了学习”,迫使其不仅仅是学习普通的文本关系。就像几周前 OpenAI 的 Jukebox(其本身也是 Scaling 在合成原始音频音乐方面的显著展示,包含极其逼真的声音/乐器)一样,GPT-3 的发布似乎几乎悄无声息,因此我将比平时更深入地探讨。

Learning to learn. In May 2020, OA released—to remarkably little interest from researchers, no blog post, no media blitz, and little public discussion beyond the snidely dismissive—the long-awaited followup to GPT-2⁠, one model to rule them all: a 117× larger 175b-parameter model with far more powerful language generation, which lets it solve a wide variety of problems from arithmetic⁠⁠1⁠ to English translation to unscrambling anagrams to SAT analogies—purely from being prompted with text examples, without any specialized training or finetuning whatsoever, merely next-word prediction training on a big Internet text corpus. This implies GPT-3’s attention mechanisms serve as “fast weights”⁠ that have “learned to learn” by training on sufficiently varied data⁠⁠2⁠, forcing it to do more than just learn ordinary textual relationships. Like OpenAI’s Jukebox⁠ just weeks ago (itself a remarkable demonstration of scaling in synthesizing _raw audio_ music complete with remarkably realistic voices/instruments), the announcement of GPT-3 appears to have sunk almost without a trace, so I will go into more depth than usual.

* GPT-3 创意小说(完整上下文):

* GPT-3 Creative Fiction⁠ (⁠full context⁠):

Flexing GPT Flexing GPT(https://gwern.net/scaling-hypothesis#flexing-gpt "Link to section: § 'Flexing GPT'")

“攻击只会变得更好。”两年前,GPT-1⁠ 凭借其有趣的预训练和可爱的“情感神经元”而引人注目。一年前,GPT-2 凭借其出色的文本生成和微调能力令人印象深刻。今年,GPT-3 令人恐惧,因为它是一个来自 2018 年初的过时架构(主要用于软件工程便利性,因为基础设施已经调试完毕),与可能的架构相比,它又小又浅⁠⁠3⁠⁠⁠4⁠,采用简单的统一架构⁠⁠5⁠,以最愚蠢的方式训练(单向预测下一个文本词元),使用单一的贫瘠模态(随机互联网 HTML 文本转储⁠⁠6⁠),数据量很小(可以放在笔记本电脑上),以愚蠢的方式采样⁠⁠7⁠,其基准性能因糟糕的提示和数据分词问题(尤其是算术和常识推理)而受损,然而,第一个版本已经表现出疯狂的运行时元学习——而且缩放曲线 _仍然_ 没有弯曲!样本也比以往任何时候都好,无论是 GPT-3 发明新的阴茎笑话⁠⁠8⁠,还是编写(大部分可用的)关于旋转数组的 JavaScript 教程。

“Attacks only get better.” 2 years ago, GPT-1⁠ was interestingly useful pretraining and adorable with its “sentiment neuron”. 1 year ago, GPT-2 was impressive with its excellent text generation & finetuning capabilities. This year, GPT-3 is scary because it’s a magnificently obsolete architecture from early 2018 (used mostly for software engineering convenience as the infrastructure has been debugged), which is small & shallow compared to what’s possible⁠⁠3⁠⁠⁠4⁠, with a simple uniform architecture⁠⁠5⁠ trained in the dumbest way possible (unidirectional prediction of next text token) on a single impoverished modality (random Internet HTML text dumps⁠⁠6⁠) on tiny data (fits on a laptop), sampled in a dumb way⁠⁠7⁠, its benchmark performance sabotaged by bad prompts & data tokenization problems (especially arithmetic & commonsense reasoning), and yet, the first version already manifests crazy runtime meta-learning—and the scaling curves _still_ are not bending! The samples are also better than ever, whether it’s GPT-3 inventing new penis jokes⁠⁠8⁠ or writing (mostly working) JavaScript tutorials about rotating arrays.

奇怪的是,这种质的飞跃似乎被标准的 NLP 基准测试在很大程度上忽略了。在 Penn Tree Bank、LAMBADA 或 WinoGrande 上报告的原始指标中,没有任何东西会让你预料到所有这些滑稽和创造性的输出;元学习结果可能会,但前提是你已经认为元学习很重要。这向我表明,GPT-3 之后的一个有用贡献是弄清楚如何对这些灵活的文本生成能力进行基准测试(可能类似于 Chollet 基于图像的抽象与推理语料库 (ARC)⁠)。

It’s odd that this qualitative leap appears to be largely missed by the standard NLP benchmarks. Nothing in the raw metrics reported on, say, Penn Tree Bank or LAMBADA or WinoGrande would lead you to expect all of this hilarious and creative output; the meta-learning results might, but only if you already thought meta-learning was important. This suggests to me that a useful post-GPT-3 contribution would be figuring out how to benchmark these sorts of flexible text generation capabilities (possibly something along the lines of Chollet’s image-based Abstraction and Reasoning Corpus (ARC)⁠).

烘焙蛋糕 Baking The Cake(https://gwern.net/scaling-hypothesis#baking-the-cake "Link to section: § 'Baking The Cake'")

并非全貌,但占很大比重。Scaling 仍在奏效。反 Scaling:因小失大。

Not the whole picture, but a big part Scaling still working Anti-scaling: penny-wise, pound-foolish

GPT 真的是 AGI 的一部分吗——还是蛋糕是个谎言?(LeCun 2019)

Is GPT actually part of AGI—or is the cake a lie? (⁠LeCun 2019⁠)

并非全貌,但占很大比重。它是否在每个任务上都达到了 SOTA?不,当然不是。但问题不在于我们能否吹毛求疵地找到它可能不奏效的任何方式,而在于是否存在它可能奏效的方式。而且有很多方式它可能表现得更好(参见“局限性”部分,仅举几例)。GPT-3 是否做了诸如操控机器人在旧金山发射激光和火箭攻击人类之类的事情?不,当然没有。它“只是”一个文本预测模型,一个文本方面的白痴天才;但我们应该记住,白痴天才距离正常人类只差一个基因突变或一点脑损伤。如果强化学习是监督学习糖霜上的樱桃,而监督学习是无监督学习蛋糕上的糖霜,那么看起来蛋糕层终于开始膨胀了。

Not the whole picture, but a big part. Does it set SOTA on every task? No, of course not. But the question is not whether we can lawyerly find any way in which it might not work, but whether there is any way which it might work⁠. And there are many ways it might work better (see the ⁠“Limitations” section⁠ for just a few). Does GPT-3 _do_ anything like steer a robot around SF shooting lasers and rockets at humans⸮ No, of course not. It is ‘just’ a text prediction model, an idiot savant of text; but an idiot savant, we should remember, is only a genetic mutation or bit of brain damage away from a normal human. If RL is the cherry on the top of the supervised learning frosting, and supervised learning is the frosting on top of the unsupervised learning cake, well, it looks like the cake layers are finally rising.

Scaling 仍在奏效。我很惊讶,因为我原本预期参数接近 1000 亿,而且我认为 CTRL/Meena/MegatronLM/T5/Turing-NLG/GPipe 的表现表明,尽管有 Scaling 论文,但 Scaling 曲线已经开始弯曲,到 1000 亿参数时,可能很难证明进一步 Scaling 的合理性。然而,在最新版本的“数据的非理性有效性”中,其中“曲线交叉”/“剪刀效应”且神经方法最终获胜(例如 Banko & Brill 2001、Brants et al 2007、Koehn & Knowles 2017),GPT-3 达到了两倍于此的参数,而 Scaling 因子没有明显变化:其 Scaling 继续大致呈对数/幂律关系,就像更小模型的情况以及预测的那样,并且它没有遇到收益实际上停止或开始需要远超可行性的增长的情况。这表明,将参数扩展到数万亿(这仍然在可用算力和预算范围内,仅需要数千个 GPU 和大约 1300 万至 1.26 亿美元 2020 年的预算,假设没有改进,当然会有改进,参见 Hernandez & Brown 2020 等)是可能且有用的,并且粗略观察图表,许多基准测试(如 Winograd 模式/WinoGrande)将在 10 万亿参数时被攻克。Scaling 的可预测性令人震惊,使得模型 Scaling 更像统计学而非人工智能。(人工智能是能按我们意愿工作但不起作用的统计学;而统计学是能起作用但不按我们意愿工作的人工智能。)

Scaling still working. I was surprised, as I had expected closer to 100b parameters, and I thought that the performance of CTRL⁠/Meena⁠/MegatronLM⁠/T5⁠/Turing-NLG⁠/GPipe⁠ suggested that, the scaling papers⁠⁠⁠9⁠ notwithstanding, the scaling curves had started to bend and by 100b, it might be hard to justify further scaling. However, in the latest version of “the unreasonable effectiveness of data”⁠ where “the curves cross”/“scissor effect” and the neural method eventually wins (eg. Banko & Brill 2001⁠, Brants et al 2007⁠, ⁠Koehn & Knowles 2017⁠), GPT-3 hits twice that without noticeable change in scaling factors: its scaling continues to be roughly logarithmic/power-law, as it was for much smaller models & as forecast, and it has not hit a regime where gains effectively halt or start to require increases vastly beyond feasibility. That suggests that it would be both possible and useful to head to trillions of parameters (which are still well within available compute & budgets, requiring merely thousands of GPUs & perhaps $13$10 2020–$126$100 2020 m budgets assuming no improvements which of course there will be, see ⁠Hernandez & Brown 2020⁠ etc.), and eyeballing the graphs, many benchmarks like the Winograd schema⁠WinoGrande⁠ would fall by 10t parameters. The predictability of scaling is striking, and makes scaling models more like statistics than AI. (AI is statistics which does what we want it to but doesn’t work; and statistics is AI which works but doesn’t do what we want.)

GPT-3:算力甚至不算多——3640 petaflop/s-day,仅是他们估计的 AlphaGo Zero 的 2 倍,1860 166ya。(历史图表由我根据“AI and Compute”,Amodei et al 2018 修改。)

GPT-3: not even that much compute—⁠3640 petaflop/s-day⁠, only 2× their estimate for AlphaGo Zero, 1860 166ya. (Historical graph modified by myself from “AI and Compute”, Amodei et al 2018⁠.)

反 Scaling:因小失大。按照机器学习标准,GPT-3 是一个极其昂贵的模型:据估计,训练它可能需要超过一只手数得过来的机器学习研究人员的年薪(约 629 万至 500 万美元 2020 年),高达 3800 万至 3000 万美元 2020 年的硬盘空间来存储模型(500–800GB),以及每 100 页输出数美分的电费(0.4 kWh)。研究人员对 Scaling 的前景感到担忧:机器学习能否负担得起成本超过 0.1 毫曼哈顿计划的项目?当然,即使它代表了 AI 能力的又一次巨大飞跃,花费高达 10 毫曼哈顿计划将 GPT-3 扩展 100 倍,以实现像在许多领域达到人类水平这样微不足道的事情,这难道不是太昂贵了吗?许多研究人员认为这样的建议是荒谬的,并驳斥了进一步扩展机器学习研究的整个想法;他们断言,他们偏好的方法(你知道的,那些不起作用的方法)将运行得更加高效,并且如果该领域转而专注于那些可以由一个贫穷的牧羊人在一台靠太阳能电池板供电的旧笔记本电脑上进行的研究,那么该领域将更具生产力。尽管如此,我认为我们可以期待进一步的 Scaling。(10 倍?不,10 倍不酷。你知道什么酷吗?100–1000 倍,在一个花哨的新超级计算机上训练。)毕竟,让某样东西在存在之后变得高效,比在存在之前更容易。

Anti-scaling: penny-wise, pound-foolish. GPT-3 is an extraordinarily expensive model by the standards of machine learning: it is estimated that training it may require the annual cost of more machine learning researchers than you can count on one hand (~$6.29$5 2020 m⁠⁠10⁠), up to $38$30 2020 of hard drive space to store the model (500–800GB), and multiple pennies of electricity per 100 pages of output (0.4 kWH). Researchers are concerned about the prospects for scaling: can ML afford to run projects which cost more than 0.1 milli-Manhattan-Projects⸮⁠⁠11⁠ Surely it would be too expensive, even if it represented another large leap in AI capabilities, to spend up to 10 milli-Manhattan-Projects to scale GPT-3 100× to a trivial thing like human-like performance in many domains⸮ Many researchers feel that such a suggestion is absurd and refutes the entire idea of scaling machine learning research further; they asseverate that their favored approaches (you know, the ones which don’t work⁠⁠12⁠) will run far more efficiently, and that the field would be more productive if it instead focused on research which can be conducted by an impoverished goatherder on an old laptop running off solar panels.⁠⁠13⁠ Nonetheless, I think we can expect further scaling. (10×? No, 10× isn’t cool. You know what’s cool? ⁠100–1000×⁠, trained on a fancy new supercomputer⁠.) It is, after all, easier to make something efficient after it exists than before.

Scaling(规模扩张) Scaling(https://gwern.net/scaling-hypothesis#scaling "Link to section: § 'Scaling'")

Scaling(规模扩张)能走多远?Scaling(规模扩张)论文表明,过去几年我们看到的飞跃,在绝对似然损失方面甚至还没有达到一半,更不用说每次额外降低损失会转化为哪些现实世界的能力。Scaling(规模扩张)曲线是清晰的;来自 Kaplan 等人 2020 年的《神经语言模型的缩放定律》:

How far will scaling go? The scaling papers suggest that the leaps we have seen over the past few years are not even half way there in terms of absolute likelihood loss, never mind what real-world capabilities each additional decrement translates into. The scaling curves are clean; from “Scaling Laws for Neural Language Models”, Kaplan et al 2020⁠:

深度学习缩放定律:算力、数据、模型参数。(图 1)

DL scaling laws: compute, data, model parameters. (⁠Figure 1⁠)

GPT-3 在此图表上代表约 10^3,为进一步降低损失留下了充足空间——尤其是考虑到外推中的不确定性:

GPT-3 represents ~10 3 on this chart, leaving plenty of room for further loss decreases—especially given the ⁠uncertainty in extrapolation⁠:

投影深度学习幂律:GPT-3 之外仍有空间。

Projecting DL power laws: still room beyond GPT-3.

果然,缩放定律在 Kaplan 等人 2020 年之后的几个数量级上继续适用于 GPT-3 模型;来自 Brown 等人 2020 年:

Lo and behold, the scaling laws continue for GPT-3 models for several orders past ⁠Kaplan et al 2020⁠; from ⁠Brown et al 2020⁠:

GPT-3 按预测继续 Scaling(规模扩张)。(注意 GPT-3 的曲线没有“反弹”,并且它只训练了约 0.5 个 epoch,见表 2.2)

GPT-3 continues to scale as predicted. (Note GPT-3’s curve has not ‘bounced’, and it trained only ~0.5 epochs, see ⁠Table 2.2⁠)

如果我们在将验证损失减半时看到如此显著的收益,但还有很长的路要走,那么当我们再次减半或减至三分之一时,还会出现什么?这到底能走多远?我们如何预测何时会出现什么?有人吗?有人吗?(另见 Meena 的困惑度与人性化聊天机器人评级、GPT-3 撰写的新闻文章按参数数量愚弄人类的概率,以及 Hendrycks 等人 2020 年关于 GPT-3 模型大小与问答的关系。)

If we see such striking gains in halving the validation loss but with so far left to go, what is left to emerge as we third or halve again? How far does this go, exactly? How do we predict what emerges when? Bueller? Bueller? (See also ⁠Meena’s perplexity vs human-ness chatbot ratings⁠, GPT-3-written news articles’ ⁠probability of fooling humans by parameter count⁠, and ⁠GPT-3 model size vs Q&A⁠ from Hendrycks et al 2020⁠.)

规模扩张的恩赐 Blessings Of Scale(https://gwern.net/scaling-hypothesis#blessings-of-scale "Link to section: § 'Blessings Of Scale'")

我们不知道如何训练神经网络。规模扩张的恩赐:稳定性 → 泛化 → 元学习

We don’t know how to train NNs Blessings of scale: stability → generalization → meta-learning

我们不知道如何训练神经网络。规模扩张的恩赐是指,对于深度学习而言,困难问题比简单问题更容易解决——随着规模变大,一切都会变得更好(这与研究的通常结果相反,通常小规模问题困难,大规模问题不可能)。神经网络/算力/数据/问题越大,它学习得越快、越好、越稳定,等等。一个在小型 n 下完全无法解决的问题,在数百万或数十亿的 n 下可能突然变得简单。“神经网络是懒惰的”:当我们推动它们超越简单的答案和廉价的捷径时,它们能做的远比我们让它们做的多。苦涩的教训是,越难越大越好。(除了 GPT-3,还可以提及半监督学习和基于模型的深度强化学习复兴的最新进展。)

We don’t know how to train NNs. The _blessings of scale_ is the observation that for deep learning, hard problems are easier to solve than easy problems—everything gets better as it gets larger (in contrast to the usual outcome in research, where small things are hard and large things impossible). The bigger the neural net/compute/data/problem, the faster it learns, the better it learns, the stabler it learns, and so on. A problem we can’t solve at all at small _n_ may suddenly become straightforward with millions or billions of _n_. “NNs are lazy”: they can do far more than we make them do when we push them beyond easy answers & cheap shortcuts. The bitter lesson is the harder and bigger, the better. (Besides GPT-3, one could mention recent progress in semi-supervised learning & the model-based DRL renaissance.)

AlphaGo Zero:“只需堆叠更多层哈哈!”

AlphaGo Zero: ‘just stack moar layers lol!’

规模扩张的恩赐:稳定性 → 泛化 → 元学习。GPT-3 受到其训练和数据的制约,但深度学习享有不合理有效的维度恩赐——仅仅在一个大数据集上训练一个大型模型就能诱导出更好的特性,如元学习,而无需任何架构内置;总的来说,在更多和更困难的任务上训练会产生更接近人类的表现、泛化能力和鲁棒性。GPT 自然语言和编程语言模型、用于图像的 iGPT/视觉 Transformer(以及某种程度上 GPT-f)表明,仅通过扩展模型和数据集而不进行任何监督,就能产生与最佳(和最复杂)替代方案相竞争的结果,使用相同的简单架构,逐渐从表面上的表面相关性过渡到更像人类的大脑活动(Schrimpf 等人 2020)和随着数据增加的语言偏见(例如 Warstadt 等人 2020)。事实上,在大规模下甚至可能不需要复杂的注意力机制,因为全连接网络——很难比它们更简单了!——在许多任务上出奇地好。人们通常使用像 Adam 这样的简单优化器来训练这样的大型模型——因为随着批量大小的增加,复杂的优化器会失去优势,而简单的优化器工作良好,并且更节省内存。OA5 不仅扩展到数百万的小批量,而且由于梯度噪声而稳定。类似 OA5 的 BigGAN 在像 JFT-300M 这样的大规模图像数据集上稳定,并受益于异常大的小批量和 VAE(长期以来在清晰图像生成方面落后于 GAN 或自回归模型),如果你让它们非常深,它们就会赶上(Child 2020,Vahdat & Kautz 2020);而像 BiT/Dojolonga 等人 2020 或 ResNeXt 或 Noisy Student 这样的分类器 CNN 会迁移并鲁棒化,具有类似人类的错误,多模态学习在更少的数据上产生更好的表示(例如 ViLBERT/VideoBERT,激发了 OA 对大型多模态模型的兴趣),RNN 可以预测视频。AlphaStar 通过数百个竞争的自对弈玩家覆盖可能的策略达到人类水平。像 MetaMimic 这样的模仿学习 DRL 在数百个任务上泛化以训练深度网络。在 StyleGAN 中,通过足够深的 w 嵌入出现解耦,有足够的参数来训练原始音频(如前述的 Jukebox),或者在关系网络/GQN/Transformer 中,有足够的样本强制分解。(另见 Hill 等人 2019/Chaplot 等人 2017/Yu 等人 2018/Lake 2019/Interactive Agents Group 2020。)在数百万次域随机化上训练 Dactyl(或人形机器人)诱导了类似的隐式元学习,在每次运行时调用中,RNN 探测其环境并将其对机器人手部控制的理解编码到其隐藏状态中;DD-PPO 通过两个数量级的扩展超越了经典机器人规划器。或者在 Procgen 或 CoinRun 中,在数百个关卡上训练会训练智能体单独解决关卡,并降低在其他关卡上的表现,但在数千个关卡上,它们开始泛化到未见过的关卡。(类似地,语言模型预训练微调在少量数据集上过拟合,但在足够多样性下显著改善。)AlphaZero 展示了真正超人的围棋,没有“妄想”,仅仅通过训练一个更大的模型在更丰富的信号上,以及无需搜索的专业级对弈——而 MuZero 则表明,仅仅端到端训练一个 RNN 来预测足够数据上的奖励,就足以使 AlphaZero 过时,并隐式地(但更好地)学习树搜索。如此等等。DM 研究员 Matthew Botvinick 在讨论他们的元强化学习工作时,惊讶地发现元学习出现了,而且无论使用哪种特定架构都会出现。

Blessings of scale: stability → generalization → meta-learning. GPT-3 is hamstrung by its training & data, but DL enjoys an unreasonably effective blessing of dimensionality⁠—just simply training a _big_ model on a _lot_ of data induces better properties like meta-learning without even the slightest bit of that architecture being built in; and in general, training on more and harder tasks creates ever more human-like performance, generalization, and robustness. The GPT natural-language & programming language models, iGPT⁠/Vision Transformer⁠ for images (and to some degree GPT-f⁠), show that simply scaling up models & datasets without any supervision produces results competitive with the best (and most complex) alternatives, using the same simple architecture, gradually passing from superficial surface correlations to more human-like brain activity (Schrimpf et al 2020⁠) and linguistic biases as data increases (eg. Warstadt et al 2020⁠). In fact, one may not even need complicated attention mechanisms at scale, as fully-connected networks—hard to get much simpler than them!—work surprisingly well⁠ for many tasks. One typically trains such large models with simple optimizers like Adam—because the complicated ones lose their advantages as batch sizes increase and the simple optimizers work fine⁠ and are more memory-efficient anyway. ⁠OA5⁠ does not just scale to, but ⁠stabilizes at⁠, minibatches of millions due to gradient noise⁠. OA5-like, BigGAN⁠ stabilizes at large-scale image datasets like JFT-300M & benefits from unusually large minibatches and VAEs (long an also-ran to GANs or autoregressive models in terms of sharp image generation) catch up if you make them very deep (Child 2020⁠, Vahdat & Kautz 2020⁠); while classifier CNNs like BiT⁠⁠⁠14⁠/Dojolonga et al 2020⁠ or ResNeXt⁠ or Noisy Student⁠ transfer &robustify⁠with⁠ human-like errors⁠⁠15⁠, multimodal learning produces better representations on fewer data (eg. ViLBERT⁠/VideoBERT⁠, motivating OA’s interest in big multimodal models⁠), and RNNs can predict videos⁠. AlphaStar⁠ reaches human-level with hundreds of competing self-players to cover possible strategies. Imitation learning DRL like MetaMimic⁠ generalizes at hundreds of tasks to train a deep net. Disentanglement emerges in StyleGAN⁠ with sufficiently deep _w_ embeddings, with enough parameters to train raw audio in the aforementioned Jukebox, or in relational networks⁠/GQN⁠/Transformers⁠ with enough samples to force factorization. (See also Hill et al 2019⁠/Chaplot et al 2017⁠/Yu et al 2018⁠/Lake 2019⁠/Interactive Agents Group 2020⁠.) Training Dactyl⁠ (or humanoid robots⁠) on millions of domain randomizations induced similar implicit meta-learning where during each runtime invocation, the RNN probes its environment and encodes its understanding of robot hand control into its hidden state; and DD-PPO⁠ outperforms classical robot planners by scaling 2 orders. Or in Procgen⁠ or CoinRun⁠, training on hundreds of levels trains agents to solve levels individually and worsens performance on other levels, but at thousands of levels, they begin to generalize to unseen levels. (Similarly, language model pretraining-finetuning⁠ overfits at small numbers of datasets but improves markedly with enough diversity.) AlphaZero⁠ demonstrated truly superhuman Go without ‘delusions’ just by training a bigger model on a richer signal & pro-level play without any search—and MuZero⁠, for that matter, demonstrated that just training an RNN end-to-end to predict a reward on enough data is enough to obsolete even AlphaZero and learn tree search implicitly (but better). And on and on. DM researcher ⁠Matthew Botvinick⁠, discussing their meta-reinforcement learning work where they were surprised to discover meta-learning emerging, and that it did so regardless of which specific architecture they used:

以 Breiman 的方式,为什么?为什么它们会迁移和泛化?为什么这些规模扩张的恩赐存在?为什么当小模型证明存在相同性能时,我们还需要训练大模型?为什么大模型不会过拟合(尽管它们可能)并且比小模型泛化得更好?整个“双重下降”到底是怎么回事?

Pace Breiman⁠, why? Why do they transfer and generalize? Why do these blessings of scale exist? Why do we need to train large models when small models provably exist with the same performance? Why do larger models not overfit (though they can⁠) and generalize better than smaller models? What’s up with the whole ‘double descent’⁠ anyway?

这些都是关于神经网络的深刻问题,并且争论激烈,但目前,我建议答案在于模型压缩/蒸馏、“彩票假说”、贝叶斯神经网络和学习表示(如电路)文献的某种混合。

These are all, ahem, deep questions about neural networks and heavily debated, but right now, I would suggest that the answer lies in some mix of the model compression/distillation, ‘lottery ticket hypothesis’⁠, Bayesian neural network⁠, and learned representation⁠ (like circuits⁠) literatures.

大模型之所以有效,是因为它们在一个极其高维的抽象空间中编码了数量惊人的子模型,代表了无数小的子模型(Orseau 等人 2020),在数据上插值,其中一个很可能很好地解决问题,从而确保问题可以被整体模型解决。它们像一个集成一样运作:即使单个大模型内部有无数过拟合的子模型,它们都会平均掉,导致对简单解的偏好。这种奥卡姆剃刀使模型偏向于简单解,这些解足够灵活,可以逐渐扩展复杂度以匹配数据。

Big models work because they encode a dizzyingly vast number of sub-models in an extremely high-dimensional abstract space, representing countless small sub-models (Orseau et al 2020⁠) interpolating over data⁠, one of which is likely to solve the problem well, and so ensures the problem is soluble by the overall model. They function as an ensemble: even though there are countless overfit sub-models inside the single big model, they all average out, leading to a preference for simple solutions. This Occam’s razor biases the model towards simple solutions which are flexible enough to gradually expand in complexity to match the data.

然而,“神经网络是懒惰的”:记忆数据片段或抓住表面特征的子模型学习最快,并且最容易在内部表示。如果模型、数据和算力不够大或不够多样化,那么在粗略训练结束时,优化只会导致一个达到低损失但错过了所需解重要部分的子模型。

However, “neural nets are lazy”: sub-models which memorize pieces of the data, or latch onto superficial features, learn quickest and are the easiest to represent internally. If the model & data & compute are not big or varied enough, the optimization, by the end of the cursory training, will have only led to a sub-model which achieves a low loss but missed important pieces of the desired solution.

另一方面,对于像 GPT-3 这样的模型,它足够强大,其子模型可以做从诗歌到算术的任何事情,并且它在如此多的数据上训练,以至于那些表面模型可能在早期表现良好,但逐渐落后于更抽象的模型;一个记忆部分数据的子模型确实比一个编码真正算术的子模型简单得多(神经网络可能用编码抽象算法如“加法”所需的空间来记忆数万个查找表条目,存储加法示例),但它不可能记忆 GPT-3 互联网规模数据集中所有算术实例(隐式或显式)。如果一个记忆子模型试图这样做,它会变得极其庞大并受到惩罚。最终,在足够的示例和更新之后,可能会出现一个相变(Viering & Loog 2021),并且准确预测数据的最简单的“算术”模型就是算术本身。然后元学习,在看到了足够多的算法实例(每个样本内略有变化,使得单独学习每个任务变得困难)之后,就是学习更通用的算法,产生比竞争对手子模型损失更低的子模型,竞争对手要么预测不好,要么膨胀得不可接受。(GPT-2-1.5b 显然太小或太浅,无法轻松集成编码元学习算法的子模型,或者可能没有在足够数据上训练足够长时间以定位元学习器模型;GPT-3 做到了。)

On the other hand, for a model like GPT-3, it is sufficiently powerful a model that its sub-models can do anything from poetry to arithmetic, and it is trained on so much data that those superficial models may do well early on, but gradually fall behind more abstract models; a sub-model which memorizes some of the data is indeed much simpler than a sub-model which encodes genuine arithmetic (a NN can probably memorize tens of thousands of lookup table entries storing examples of addition in the space it would take to encode an abstract algorithm like ‘addition’), but it can’t possibly memorize _all_ the instances of arithmetic (implicit or explicit) in GPT-3’s Internet-scale dataset. If a memorizing sub-model tried to do so, it would become extremely large and penalized. Eventually, after enough examples and enough updates, there may be a phase transition (⁠Viering & Loog 2021⁠), and the simplest ‘arithmetic’ model which accurately predicts the data just _is_ arithmetic. And then the meta-learning, after seeing enough instances of algorithms which vary slightly within each sample, making it hard to learn each task separately, just _is_ learning of more generic algorithms, yielding sub-models which achieve lower loss than the rival sub-models, which either fail to predict well or bloat unacceptably. (GPT-2-1.5b apparently was too small or shallow to ensemble easily over sub-models encoding meta-learning algorithms, or perhaps not trained long enough on enough data to locate the meta-learner models; GPT-3 was.)

因此,模型越大越好,只要有足够的数据和算力将其推过容易的子模型,进入表达理想特性的子模型,如泛化、将感知分解为有意义的潜在维度、基于描述进行元学习、学习因果推理和逻辑等。如果成分存在,它就会发生。

So, the larger the model, the better, if there is enough data & compute to push it past the easy convenient sub-models and into the sub-models which express desirable traits like generalizing, factorizing perception into meaningful latent dimensions, meta-learning tasks based on descriptions, learning causal reasoning & logic, and so on. If the ingredients are there, it’s going to happen.

“规模扩张的恩赐”的反向链接(18):

Backlinks (18)⁠ for ⁠“Blessings Of Scale”⁠:

* 迈向基准测试 LLM 多样性与创造力(完整上下文):

* Towards Benchmarking LLM Diversity & Creativity⁠ (⁠full context⁠):

* 绝对单元神经网络:用于所有事物的基于回归的 MLP(完整上下文):

* Absolute Unit NNs: Regression-Based MLPs for Everything⁠ (⁠full context⁠):

* 扩展 MLP:归纳偏好的故事:

* Scaling MLPs: A Tale of Inductive Bias⁠:

* 用于上传的模块化大脑 AUNN(完整上下文):

* Modular Brain AUNNs for Uploads⁠ (⁠full context⁠):

* RL 智能体的自由游戏期(完整上下文):

* Free-Play Periods for RL Agents⁠ (⁠full context⁠):

* GAN 没有失败,它们被抛弃了(完整上下文):

* GANs Didn’t Fail, They Were Abandoned⁠ (⁠full context⁠):

* 思维链提示引发大型语言模型中的推理:

* Chain-of-Thought Prompting Elicits Reasoning in Large Language Models⁠:

* 顿悟:超越小算法数据集上的过拟合泛化:

* Grokking: Generalization Beyond Overfitting On Small Algorithmic Datasets⁠:

* GPT-2 偏好学习用于音乐生成(完整上下文):

* GPT-2 Preference Learning for Music Generation⁠ (⁠full context⁠):

* ‘神经网络稀疏性’目录(完整上下文):

* ‘NN sparsity’ directory⁠ (⁠full context⁠):

* ‘AI 扩展’目录(完整上下文):

* ‘AI scaling’ directory⁠ (⁠full context⁠):

* ARPA 和 SCI:冲浪 AI(完整上下文):

* ARPA and SCI: Surfing AI⁠ (⁠full context⁠):

Scaling 假设 Scaling Hypothesis(https://gwern.net/scaling-hypothesis#scaling-hypothesis "Link to section: § 'Scaling Hypothesis'")

强 _Scaling 假设_ 认为,一旦我们找到像自注意力或卷积这样的可扩展架构,它们可以像大脑一样相当均匀地应用(例如,“大脑作为通用学习机器”或 Hawkins),我们就可以简单地训练越来越大的神经网络,而越来越复杂的行为将作为优化所有任务和数据的最简单方式自然涌现。更强大的神经网络‘只是’放大了的弱神经网络,就像人类大脑看起来像是放大了的灵长类动物大脑一样。

The strong _scaling hypothesis_ is that, once we find a scalable architecture like self-attention or convolutions, which like the brain can be applied fairly uniformly (eg. “The Brain as a Universal Learning Machine”⁠ or Hawkins), we can simply train ever larger NNs and ever more sophisticated behavior will emerge naturally as the easiest way to optimize for all the tasks & data. More powerful NNs are ‘just’ scaled-up weak NNs, in much the same way that human brains look much like scaled-up primate brains⁠.

虽然我在 2004-2006 年、2010 年、2016 年首次对 AI 产生兴趣时(当时 AI 还陷在 hopelessly narrow tools 的低谷,像 2028 这样的日期似乎遥不可及)对 Scaling 假设的支持者高度怀疑,认为这带有数字命理学和“如果你建造它,他们就会来”的逻辑(当时我们当然没有可以随意投入算力的通用算法),但在 2020 年,我不得不承认,我错了,他们是对的。我们建造了算力,而算法 _确实_ 出现了,并且自 2010 年、2016 年以来,Scaling 假设每年都显得越来越合理。

While I was highly skeptical of scaling hypothesis advocates when I first became interested in AI 2004–6 2010 16ya (back when AI was stuck in the doldrums of hopelessly narrow tools and dates like 2028 seemed impossibly far away), which smacked of numerology and “if you build it they will come” logic (at the time, we certainly didn’t have general algorithms that you could just throw compute at), in 2020, I have to admit, I was wrong and they were right. We built the compute, and the algorithms _did_ come, and the scaling hypothesis has only looked more and more plausible every year since 2010 16ya.

反向链接 (2)⁠ 用于 ⁠“Scaling 假设”⁠:

Backlinks (2)⁠ for ⁠“Scaling Hypothesis”⁠:

预训练为何有效? Why Does Pretraining Work?(https://gwern.net/scaling-hypothesis#why-does-pretraining-work "Link to section: § 'Why Does Pretraining Work?'")

最后几比特是最深刻的怀疑理由。

The last bits are deepest Reasons for doubt

预训练论点大致如下:

The pretraining thesis goes something like this:

“图 1:NLP 研究通过三个不同时代或曲线的设想演变”(假设的 S 曲线与自然语言建模的进展;来自 Cambria & White 2014)

“Figure 1: Envisioned evolution of NLP research through three different eras or curves” (the hypothetical S-curves & progress in natural language modeling; from Cambria & White 2014⁠)

可以说,人类是 AI 的蓝细菌:我们不断排放大量结构化数据,这些数据隐含地依赖于逻辑、因果、物体恒存、历史——所有那些好东西。所有这些都隐含并编码在我们的文字、视频和“数据废气”中。一个学习预测的模型必须学会理解所有这些才能获得最佳性能;当它预测那些仅仅是统计模式匹配的简单事物时,剩下的就是困难的事物。AI 批评者常说,自动驾驶或自然语言等任务的长期尾场景只能通过真正的泛化和推理来解决;因此,如果模型解决了长尾,它们必须学会泛化和推理。

Humans, one might say, are the cyanobacteria of AI⁠: we constantly emit large amounts of structured data, which implicitly rely on logic, causality, object permanence, history—all of that good stuff. All of that is implicit and encoded into our writings and videos and ‘data exhaust’. A model learning to predict must learn to understand all of that to get the best performance; as it predicts the easy things which are mere statistical pattern-matching, what’s left are the hard things. AI critics often say that the long tail of scenarios for tasks like self-driving cars or natural language can only be solved by true generalization & reasoning; it follows then that if models solve the long tail, they must learn to generalize & reason.

在训练早期,模型学习最粗糙的层次:某些字母如'e'比'z'更常见,每 5 个字符左右有一个空格,等等。它从预测均匀分布的字节转变为看起来像 Base-60 编码——字母数字乱码。尽管这很粗糙,但足以取得相当大的绝对进步:随机预测器需要 8 比特来“预测”一个字节/字符,但仅仅通过匹配字母和空格频率,它就能将误差几乎减半至约 5 比特。⁠⁠16⁠ 由于它从每个字符中学到很多,而且学习的频率很简单,这可以发生得如此之快,以至于如果不频繁记录样本,人们甚至可能观察不到这一改进。

Early on in training, a model learns the crudest levels: that some letters like ‘e’ are more frequent than others like ‘z’, that every 5 characters or so there is a space, and so on. It goes from predicted uniformly-distributed bytes to what looks like Base-60 encoding—alphanumeric gibberish. As crude as this may be, it’s enough to make quite a bit of absolute progress: a random predictor needs 8 bits to ‘predict’ a byte/character, but just by at least matching letter and space frequencies, it can almost halve its error to around 5 bits.⁠⁠16⁠ Because it is learning so much from every character, and because the learned frequencies are simple, it can happen so fast that if one is not logging samples frequently, one might not even observe the improvement.

随着训练进行,任务变得更加困难。现在它开始学习哪些单词实际存在,哪些不存在。它不知道任何意义,但至少当被要求预测一个单词的后半部分时,它能在一定程度上做到,从而节省几个比特。这需要一段时间,因为任何特定实例只会偶尔出现:一个单词可能不会出现在十几个样本中,而且有成千上万个单词要学习。经过更多努力,它学会了标点、复数、所有格都是存在的东西。综合起来,它可能又取得了进步,一直降到每个字符 3-4 比特的误差!(虽然进步快得令人欣慰,但仍然是乱码,毫无疑问:一个样本可能拼写正确,但毫无意义。)

As training progresses, the task becomes more difficult. Now it begins to learn what words actually exist and do not exist. It doesn’t know anything about meaning, but at least now when it’s asked to predict the second half of a word, it can actually do that to some degree, saving it a few more bits. This takes a while because any specific instance will show up only occasionally: a word may not appear in a dozen samples, and there are many thousands of words to learn. With some more work, it has learned that punctuation, pluralization, possessives are all things that exist. Put that together, and it may have progressed again, all the way down to 3–4 bits error per character! (While the progress is gratifyingly fast, it’s still all gibberish, though, makes no mistake: a sample may be spelled correctly, but it doesn’t make even a bit of sense.)

但是一旦模型学会了良好的英语词汇和正确的格式/拼写,接下来呢?在单词内部预测方面已经没有多少油水了。接下来是捕捉单词之间的关联。哪些单词倾向于先出现?哪些单词“聚类”并经常在彼此附近使用?航海术语在海事故事中经常一起使用,圣经段落、美国历史维基百科文章也是如此。如果最后一个词是“Jefferson”,那么“Washington”可能不远,它应该对冲赌注预测下一个字符是'W',然后如果出现,就全力押注“ashington”。这种词袋方法仍然预测得很差,但现在我们可能降到每个字符<3 比特。

But once a model has learned a good English vocabulary and correct formatting/spelling, what’s next? There’s not much juice left in predicting within-words. The next thing is picking up associations among words. What words tend to come first? What words ‘cluster’ and are often used nearby each other? Nautical terms tend to get used a lot with each other in sea stories, and likewise Bible passages, or American history Wikipedia article, and so on. If the word “Jefferson” is the last word, then “Washington” may not be far away, and it should hedge its bets on predicting that ‘W’ is the next character, and then if it shows up, go all-in on “ashington”. Such bag-of-words approaches still predict badly, but now we’re down to perhaps <3 bits per character.

接下来呢?它会停在那里吗?如果数据足够多,并且早期学习英语词汇等没有耗尽模型的学习能力,就不会。逐渐地,其他词如“President”或“general”或“after”开始向模型展示微妙的关联:“Jefferson was President after…”通过许多这样的段落,单词“after”开始用于预测下一个词,然后其用途可以扩大。

What next? Does it stop there? Not if there is enough data and the earlier stuff like learning English vocab doesn’t hem the model in by using up its learning ability. Gradually, other words like “President” or “general” or “after” begin to show the model subtle correlations: “Jefferson was President after…” With many such passages, the word “after” begins to serve a use in predicting the next word, and then the use can be broadened.

到这时,损失可能约为 2 比特:每额外降低 0.1 比特都付出更昂贵的代价,需要更多时间。然而,现在句子开始有意义了。像“Jefferson was President after Washington”这样的句子确实有意义(如果偶尔我们采样到“Washington was President after Jefferson”,嗯,你对这样一个未收敛的模型还能期待什么呢)。刺眼的错误会立即将我们从任何关于模型理解的幻觉中惊醒,于是训练继续。(大约在这里,马尔可夫链和 n-gram 模型开始落后;它们可以记住训练语料库中越来越大的块,但无法解决越来越关键的句法任务,如平衡括号或引号,更不用说开始从句法上升到语义了。)

By this point, the loss is perhaps 2 bits: every additional 0.1 bit decrease comes at a steeper cost and takes more time. However, now the sentences have started to make sense. A sentence like “Jefferson was President after Washington” does in fact mean something (and if occasionally we sample “Washington was President after Jefferson”, well, what do you expect from such an un-converged model). Jarring errors will immediately jostle us out of any illusion about the model’s understanding, and so training continues. (Around here, Markov chain &_n_-gram models start to fall behind; they can memorize increasingly large chunks of the training corpus, but they can’t solve increasingly critical syntactic tasks like balancing parentheses or quotes, much less start to ascend from syntax to semantics.)

现在训练变得困难。必须建模语言中更微妙的方面,例如保持代词一致。这之所以困难,部分是因为模型的错误变得罕见,而且相关的文本片段越来越遥远和“长程”。随着它取得进展,误差的绝对大小急剧缩小。考虑将名字与性别代词关联的情况:“Janelle ate some ice cream, because he likes sweet things like ice cream”和“Janelle ate some ice cream, because she likes sweet things like ice cream”之间的差异是人类无法忽视的,然而,这只是一个字母的差异。如果我们比较两个模型,一个完全不懂性别代词,纯粹随机猜测'he'/'she',另一个完全理解并总是猜测'she',第二个模型将获得更低的平均误差,仅略低于每个字符 0.02 比特!

Now training is hard. Even subtler aspects of language must be modeled, such as keeping pronouns consistent. This is hard in part because the model’s errors are becoming rare, and because the relevant pieces of text are increasingly distant and ‘long-range’. As it makes progress, the absolute size of errors shrinks dramatically. Consider the case of associating names with gender pronouns: the difference between “Janelle ate some ice cream, because he likes sweet things like ice cream” and “Janelle ate some ice cream, because she likes sweet things like ice cream” is one no human could fail to notice, and yet, it is a difference of a single letter. If we compared two models, one of which didn’t understand gender pronouns at all and guessed ‘he’/‘she’ purely at random, and one which understood them perfectly and always guessed ‘she’, the second model would attain a lower average error of barely <0.02 bits per character!

尽管如此,随着训练继续,这些问题以及更多问题,如模仿体裁,得到解决,最终在损失为 1-2 时(一个小型字符 RNN 可能在莎士比亚或某些古腾堡计划电子书等小型语料库上收敛),我们将最终得到听起来像人类的样本——至少,对于几个句子。这些最终样本可能暂时说服我们,但除了重复循环等问题外,即使有好的样本,误差也会累积:一个样本会声称某人“活着”,然后 10 个句子后使用“死了”,或者它会离题到不相关的论点而不是预期的下一个论点,或者某人会做出物理上不可能的事情,或者它可能只是继续一段时间而似乎没有取得任何进展。

Nevertheless, as training continues, these problems and more, like imitating genres, get solved, and eventually at a loss of 1–2 (where a small char-RNN might converge on a small corpus like Shakespeare or some Project Gutenberg ebooks), we will finally get samples that sound human—at least, for a few sentences. These final samples may convince us briefly, but, aside from issues like repetition loops, even with good samples, the errors accumulate: a sample will state that someone is “alive” and then 10 sentences later, use the word “dead”, or it will digress into an irrelevant argument instead of the expected next argument, or someone will do something physically improbable, or it may just continue for a while without seeming to _get_ anywhere.

所有这些误差都远小于每个字符 0.02 比特;我们现在讨论的不是百分之一比特,而是小于万分之一比特。

All of these errors are far less than <0.02 bits per character; we are now talking not hundredths of bits per characters but less than ten-thousandths.

预训练论点认为这可以走得更远:我们可以直接将这种性能与人类执行相同目标任务的表现进行比较,人类可以达到约每个字符 0.7 比特。那缺失的>0.4 比特中有什么?

The pretraining thesis argues that this can go even further: we can compare this performance directly with humans doing the same objective task, who can achieve closer to ⁠0.7 bits per character⁠. What is in that missing >0.4?

“是啊,但聪明不仅仅是知道压缩方案!” “不,就是!” “糟糕——他知道秘密了!!”

“Yeah, but there’s more to being smart than knowing compression schemes!” “No there’s not!” “Shoot—he knows the secret!!”

嗯——一切!模型错过的一切。虽然一开始胡言乱语就足够了,但最终,它需要能够推理出最困难的文本场景,这些场景需要因果或常识推理。每一次模型预测冰淇淋放在冰箱里会“融化”而不是“冻结”,每一次模型无法分清一个人是活着还是死了,每一次模型选择一个无助于最终构建“文章”结论的词,每一次它缺乏心智理论来压缩描述十几个人在晚餐时玩弄权术、争权夺利的新颖场景,每一次使用逻辑、抽象、指令或问答时模型困惑并需要更多比特来掩盖其错误,而人类会思考、理解和预测。对于语言模型,真理就是那些能持续良好预测的东西——因为真理是唯一的,而错误是多样的。这些认知突破中的每一个都允许对少量相关文本进行略微更好的预测;只有真正的理解才能满足理想预测。

Well—_everything_! Everything that the model misses. While just babbling random words was good enough at the beginning, at the end, it needs to be able to reason our way through the most difficult textual scenarios requiring causality or commonsense reasoning. Every error where the model predicts that ice cream put in a freezer will “melt” rather than “freeze”, every case where the model can’t keep straight whether a person is alive or dead, every time that the model chooses a word that doesn’t help build somehow towards the ultimate conclusion of an ‘essay’, every time that it lacks the theory of mind to compress novel scenes describing the Machiavellian scheming of a dozen individuals at dinner jockeying for power as they talk, every use of logic or abstraction or instructions or Q&A where the model is befuddled and needs more bits to cover up for its mistake where a human would think, understand, and predict. For a language model, the truth is that which keeps on predicting well—because truth is one and error many. Each of these cognitive breakthroughs allows ever so slightly better prediction of a few relevant texts; nothing less than true understanding will suffice for ideal prediction.

如果我们训练一个模型达到那个<0.7 的损失,它能够预测与人类无法区分的文本,无论是在对话中、被问及冰淇淋、接受 SAT 类比测试还是接受数学辅导,如果对于每个字符串,模型在预测下一个字符方面做得和你一样好,我们怎么能说它没有真正理解一切呢?(至少,根据定义,我们可以用模型取代任何文本写作工作中的人类!)

If we trained a model which reached that loss of <0.7, which could predict text indistinguishable from a human, whether in a dialogue or quizzed about ice cream or being tested on SAT analogies or tutored in mathematics, if for every string the model did just as good a job of predicting the next character as you could do, how could we say that it doesn’t _truly_ understand everything? (If nothing else, we could, by definition, replace humans in any kind of text-writing job!)

最后几比特是最深刻的。这里的含义是,最后几比特是最有价值的比特,它们需要我们认为的智能中最重要的部分。Collobert 等人 2011:

The last bits are deepest. The implication here is that the final few bits are the most valuable bits, which require the most of what we think of as intelligence. Collobert et al 2011⁠:

这里一个有用的类比可能是我们的行动:在大多数情况下,所有人类执行行动的能力都一样好。我们都拿起茶杯而不掉落,可以抬起腿走下数千级台阶而一次也不摔倒。对于日常行动(构成语料库大部分的那种),任何智力水平的人都能通过足够的练习和反馈做得很好,学习单独的算法来孤立地解决每类问题,而且解决得非常好。⁠⁠17⁠ 同时,对于罕见问题,可能实例太少,除了记住答案外别无他法。在频谱中间是那些与其他问题相似但又不那么相似的问题;这些是奖励灵活元学习和泛化的问题,并且可能需要许多中间问题来引出这些能力(“神经网络是懒惰的”)。

A helpful analogy here might be our actions: for the most part, all humans execute actions equally well. We all pick up a tea mug without dropping, and can lift our legs to walk down thousands of steps without falling even once. For everyday actions (the sort which make up most of a corpus), anybody, of any intelligence, can get enough practice & feedback to do them quite well, learning individual algorithms to solve each class of problems extremely well, in isolation.⁠⁠17⁠ Meanwhile for rare problems, there may be too few instances to do any better than memorize the answer. In the middle of the spectrum are problems which are similar but not _too_ similar to other problems; these are the sorts of problem which reward flexible meta-learning and generalization, and many intermediate problems may be necessary to elicit those capabilities⁠ (“neural nets are lazy”).

个体之间的差异在于当他们开始遇到长尾的新奇选择、罕见选择、需要几秒钟但影响一生的选择、我们永远得不到反馈的选择(如死后)时。一个人只需要在一生数百万个离散决策中做出一个糟糕的决定,就可能进监狱或死亡。决策质量的一个微小绝对平均改进,如果是在那些决策中,可能远比其数量所指示的重要得多,并让我们直观理解为什么最后几比特是最困难/最深刻的。(为什么人类有如此大的大脑,而像黑猩猩这样的动物似乎以一小部分代价就能同样好地完成许多日常活动?为什么语言值得?也许是因为这些考虑。我们在填写人寿保险文件时可能最像人类。)

Where individuals differ is when they start running into the long tail of novel choices, rare choices, choices that take seconds but unfold over a lifetime, choices where we will never get any feedback (like after our death). One only has to make a single bad decision, out of a lifetime of millions of discrete decisions, to wind up in jail or dead. A small absolute average improvement in decision quality, if it is in _those_ decisions, may be far more important than its quantity indicates, and give us some intuition for why those last bits are the hardest/deepest. (Why do humans have such large brains, when animals like chimpanzees do so many ordinary activities seemingly as well with a fraction of the expense? Why is language worthwhile? Perhaps because of considerations like these. We may be at our most human while filling out the paperwork for life insurance.)

怀疑的理由。预训练论点虽然在逻辑上无懈可击——一个模型怎么可能在不理解的情况下解决所有可能的刁钻问题,而只是猜测?——但从未让我觉得有说服力,这是一个既不接受反驳也不接受信服的论点。它感觉太像魔术把戏:“这里有一些信息论,这里有一个人类基准,这里是我们如何将所有任务编码为序列预测问题,嘿,变——智能!”有很多算法在某种意义上是图灵完备或“通用”的;有很多算法如 AIXI 在某种理论意义上解决了 AI(Schmidhuber 及其公司有许多这样可爱的算法,如“所有问题的最快可能算法”,但有一个小问题:需要比宇宙还大的计算机的某些常数因子)。

Reasons for doubt. The pretraining thesis, while logically impeccable—how is a model supposed to solve all possible trick questions without understanding, just _guessing_?—never struck me as convincing, an argument admitting neither confutation nor conviction. It feels too much like a magic trick: “here’s some information theory, here’s a human benchmark, here’s how we can encode all tasks as a sequence prediction problem, hey presto—Intelligence!” There are lots of algorithms which are Turing-complete or ‘universal’ in some sense; there are lots of algorithms like AIXI which solve AI in some theoretical sense (Schmidhuber & company have many of these cute algorithms such as ‘the fastest possible algorithm for all problems’, with the minor catch of some constant factors which require computers bigger than the universe).

为什么认为预训练或序列建模不是其中之一?当然,如果模型获得了足够低的损失,它就必须是智能的,但你如何证明这在实践中会发生?(训练字符 RNN 很有趣,但它们并没有彻底改变深度学习。)它可能需要比现存更多的文本,无数 PB 的数据,才能让那些微妙因素如逻辑推理在噪声和干扰中提供足够的训练信号来训练模型。或者你的模型太小,只能吸收简单的表面信号,你需要将它们缩放 100 个数量级才能工作,因为缩放曲线不配合。或者你的模型从根本上是有缺陷的,像抽象这样的东西需要完全不同的架构才能工作,无论你做什么,你当前的模型都会在糟糕的性能上饱和。或者它会训练,但会花费所有时间试图改进表面建模,吸收越来越多的字面数据和事实,而从未按计划上升到更高的认知层面。或者……

Why think pretraining or sequence modeling is not another one of them? Sure, _if_ the model got a low enough loss, it’d have to be intelligent, but how could you prove that would happen in practice? (Training char-RNNs was fun, but they hadn’t exactly revolutionized deep learning.) It might require more text than exists, countless petabytes of data for all of those subtle factors like logical reasoning to represent enough training signal, amidst all the noise and distractors, to train a model. Or maybe your models are too small to do more than absorb the simple surface-level signals, and you would have to scale them 100 orders of magnitude for it to work, because the scaling curves didn’t cooperate. Or maybe your models are fundamentally broken, and stuff like abstraction require an entirely different architecture to work at all, and whatever you do, your current models will saturate at poor performance. Or it’ll train, but it’ll spend all its time trying to improve the surface-level modeling, absorbing more and more literal data and facts without ever ascending to the higher planes of cognition as planned. Or…

但显然,它本来就会工作得很好。即使是 RNN 也可能工作——Transformer 很好,但它们似乎主要是关于效率。⁠⁠19⁠(训练大型 RNN 要昂贵得多,并且在多个节点上进行 BPTT 在工程上要困难得多。)它只是需要比任何人愿意冒险投入的更多的算力和数据,直到少数真正相信的人能够获得几百万美元的算力。

But apparently, it would’ve worked fine. Even RNNs probably would’ve worked—Transformers are nice, but they seem mostly be about efficiency.⁠⁠19⁠ (Training large RNNs is much more expensive, and doing BPTT over multiple nodes is much harder engineering-wise.) It just required more compute & data than anyone was willing to risk on it until a few true-believers were able to get their hands on a few million dollars of compute.

* 问:有没有人定量预测过这会在何时发生?

* Q:Did anyone predict, quantitatively, that this would happen where it did?

* 问:未来更大规模的模型会学到什么?

* Q:What would future scaled-up models learn?

GPT-2-1.5b 的交叉熵 WebText 验证损失约为 3.3(基于图 4 中约 10 的困惑度,log2(10)=3.32)。GPT-3 根据 Brown 等人 2020 并使用缩放公式(2.57 × (3.64 × 10^3)^(-0.048))将该损失减半至约 1.73。对于一个假设的 GPT-4,如果缩放曲线在交叉并遇到更严重的收益递减之前再持续大约 3 个数量级的算力(100-1000 倍),交叉熵损失将降至约 1.24(2.57 × (3.64 × (10^3 × 10^3))^(-0.048))。

GPT-2-1.5b had a cross-entropy WebText validation loss of ~3.3 (based on the perplexity of ~10 in ⁠Figure 4⁠, and log 2(10) = 3.32). GPT-3 halved that loss to ~1.73 judging from ⁠Brown et al 2020⁠ and using the scaling formula (2.57 × (3.64 × 10 3)−0.048). For a hypothetical GPT-4, if the scaling curve continues for another 3 orders or so of compute (100–1000×) before crossing over and hitting harder diminishing returns, the cross-entropy loss will drop to ~1.24 (2.57 × (3.64 × (10 3 × 10 3))−0.048).

如果 GPT-3 通过从 GPT-2 的水平降低绝对损失约 50%获得了如此多的元学习和世界知识,那么相对于 GPT-3 再改进约 30%会获得什么能力?(将损失降低那么多仍然不会达到人类水平,据我所知。⁠⁠20⁠)降到≤1,也许通过更宽的上下文窗口或循环,会获得什么?

If GPT-3 gained so much meta-learning and world knowledge by dropping its absolute loss ~50% when starting from GPT-2’s level, what capabilities would another ~30% improvement over GPT-3 gain? (Cutting the loss that much would still not reach human-level, as far as I can tell.⁠⁠20⁠) What would a drop to ≤1, perhaps using wider context windows or recurrency, gain?

指向“预训练为何有效?”的反向链接(2):

Backlinks (2)⁠ for ⁠“Why Does Pretraining Work?”⁠:

展望 Prospects(https://gwern.net/scaling-hypothesis#prospects "Link to section: § 'Prospects'")

我们可以期待未来深度学习工作带来什么?GPT-3 是否会引发一场军备竞赛,以至于我们很快就能平淡地讨论那些现在看来荒谬离奇的方案,比如一个双向多模态 Transformer,其规模扩大 100 倍,训练数据增加 100 倍(视频/文本/PDF 作为图像/照片/机器人),并辅以监督学习,作为类似 MuZero 的学习+规划深度强化学习智能体的骨干,同时在数千个任务(如编程)上运行?[大致上,是的。——编者 2025-10-19]

What can we expect from future DL work? Will GPT-3 kickstart an arms race where soon we will be discussing, blasé, what would seem now like ludicrously farfetched schemes like bidirectional multimodal Transformer 100× the size trained on 100× the data (video/text/PDFs-as-images/photo/robotics) with supplementary supervised learning as the backbone of a MuZero-like learning+planning DRL agent running on thousands of tasks (such as coding) simultaneously? [Roughly, yes. —Editor 2025-10-19]

硬件过剩的存在意味着这里的限制因素与其说是硬件,不如说是人力:会有组织将 GPT-3 视为斯普特尼克时刻,并积极投资于 Scaling 项目吗?DeepMind 或 Google Brain 的 TPU Pod 中是否正在酝酿着相当于 GPT-4 的东西?他们并不愚蠢,他们有硬件,有预算,也有人才。

The existence of the hardware overhang⁠ implies that the limiting factor here is less hardware than human: will any organization treat GPT-3 as a Sputnik moment and invest aggressively in scaling programs? Is there a GPT-4-equivalent brewing away inside DeepMind or Google Brain’s TPU pods now? They aren’t stupid, they have the hardware, they have the budgets, they have the people.

但我认为他们缺乏远见。据我所知:他们没有这样的东西,因为 Google Brain 和 DeepMind 并不像 Sutskever、Amodei 和 OpenAI 的其他人那样相信 Scaling 假设。只需浏览机器学习推特,就能看到对 Scaling 假设的蔑视。(GPT-3 发布已过四分之一年,你能说出一个像 17B 的 Turing-NLG 那样大的密集模型吗——更不用说比 GPT-3 更大的了?)

But I think they lack a vision. As far as I can tell: they do not have any such thing, because Google Brain & DeepMind do not believe in the scaling hypothesis the way that Sutskever, Amodei and others at OA do. Just read through machine learning Twitter to see the disdain for the scaling hypothesis. (A quarter year on from GPT-3 and counting, can you name a single dense model as large as the 17b Turing-NLG—never mind larger than GPT-3?)

Google Brain 过于务实和短视,不会涉足如此深奥且昂贵的投机,尽管 Quoc V. Le 的团队偶尔会给你惊喜。他们会涉足混合专家模型,如 GShard,但主要是因为他们期望能够将其或类似的东西部署到 Google 翻译的生产环境中。

Google Brain is entirely too practical and short-term focused to dabble in such esoteric & expensive speculation, although Quoc V. Le’s group occasionally surprises you. They’ll dabble in mixture-of-expert models⁠ like GShard⁠, but mostly because they expect to be likely to be able to deploy it or something like it to production in Google Translate.⁠⁠23⁠

为什么 DeepMind 没有做出 GPT-3?DeepMind 持有我们可称之为“弱 Scaling 假设”的观点:他们认为 AGI 需要我们“找到正确的算法”,有效地逐个模块复制哺乳动物的大脑,并且虽然这些模块按当代标准将极其庞大和昂贵(这就是为什么算力很重要,它给了我们“一个更强大的工具来寻找正确的算法”),但它们仍然需要逐个发明和微调,直到最终组装完成之前几乎没有风险或意外。然而,每个模块本身可以 Scaling:不存在神奇的智能腺体或量子胡话在人类与黑猩猩或啮齿动物之间划出一条清晰的分界线。(尽管我们人类过度欣赏自己的语言或逻辑能力,但这些只是基本大脑上的相对次要的点缀——每个有机体都解决相同的基本问题,如探索、长期记忆、学习世界模型、将奖励与特定动作关联、元学习等。)因此,一旦你有了老鼠级别的 AGI,人类级别的 AGI 只是规模更大而已。(而且老鼠更容易做实验。)这就是为什么你会看到像 Agent57 这样的 DeepMind 装置,把厨房水槽扔到墙上看看什么能粘住,以及为什么他们如此强调神经科学作为逆向工程大脑的灵感和交叉融合。(另见 Sam Altman 的播客采访评论,关于 OpenAI 相对于拥有更多算力的未命名竞争对手的优势,是因为缺乏算力使他们保持“小而专注”——“当然”像初创公司的方法。)当有人似乎提出了一个可扩展的架构来解决一个难题,比如 AlphaZero 或 AlphaStar,他们愿意加大投入使其扩展,但除此之外,在 ALE 和 DMLab-30 上的渐进式改进是游戏计划。他们十年来一直在啃食大脑的碎片,如果一切顺利,可能还需要十年或二十年的稳步啃食。因为他们锁定了如此多的人才,拥有如此多的专有代码,并相信所有这些对任何试图复制复杂大脑的竞争对手来说都是一条重要的护城河,所以他们相当从容。你不会看到 DeepMind 在任何登月计划上“押上公司”;谷歌的现金流不会消失(DeepMind 的预算也是),稳扎稳打赢得比赛。

Why didn’t DeepMind do GPT-3?⁠ DeepMind⁠⁠24⁠ holds what we might call the “weak scaling hypothesis”: they believe that AGI will require us to “find the right algorithms” effectively replicating a mammalian brain module by module, and that while these modules will be extremely large & expensive by contemporary standards (which is why compute is important, to give us “a more powerful tool with which to hunt for the right algorithms”), they still need to be invented & finetuned piece by piece, with little risk or surprise until the final assembly. Each piece, however, itself can scale: there’s no magical intelligence gland or quantum woo which creates a bright line between humans and, say, chimpanzees or rodents. (As much as we humans extravagantly admire our own capabilities like language or logic, those are relatively minor flourishes on the basic brain—each organism solves the same basic problems, like exploration, long-term memory, learning world-models, associating rewards with specific actions, meta-learning, etc.) As such, once you have a rat-level AGI, a human-level AGI is just more so. (And rats are a lot easier to experiment on.) That is how you get DM contraptions like Agent57⁠ which throw the kitchen sink at the wall to see what sticks, and why they place such emphasis on neuroscience as inspiration and cross-fertilization for reverse-engineering the brain. (See also Sam Altman’s ⁠podcast interview comments⁠ on OA’s advantage vs unnamed rivals with more compute is because the lack of compute makes them stay “small and focused”—“for sure” like a startup approach.) When someone seems to have come up with a scalable architecture for cracking a hard problem, like AlphaZero or AlphaStar, they are willing to pour on the gas to make it scale, but otherwise, incremental refinement on ALE and then DMLab-30⁠ is the game plan. They have been biting off and chewing pieces of the brain for a decade, and it’ll probably take another decade or two of steady chewing if all goes well. Because they have locked up so much talent and have so much proprietary code and believe all of that is a major moat to any competitor trying to replicate the complicated brain, they are fairly easygoing. You will not see DM ‘bet the company’ on any moonshot; Google’s cashflow isn’t going anywhere (and ⁠DM’s budget⁠), and slow and steady wins the race.

除此之外,大多数其他研究实验室,如特斯拉或 FAIR,要么无关紧要,要么不感兴趣。中国的人工智能公司是一个问号:跨越语言障碍,我似乎看到了对 AGI 的兴趣,而西方那种条件反射式的反对很少,像百度这样的公司偶尔会发布重要的研究(如早期的 Scaling 论文 Hestness et al 2017),但总体而言,中国的人工智能可能被高估了,他们似乎遭受了一种荷兰病——监控技术和狭窄电子商务领域的资金如此充裕,以至于其他领域被忽视。

Going beyond that, most other research labs like Tesla or FAIR are irrelevant and uninterested. Chinese AI companies are a question mark: past the language barrier, I seem to discern interest in AGI & little of the reflexive Western opposition, and companies like Baidu occasionally release important research (such as the early scaling paper Hestness et al 2017⁠), but overall, Chinese AI may be overestimated, and they seem to suffer from a kind of Dutch disease—funding for surveillance technology, and for narrow e-commerce niches, is so plentiful that other areas are neglected.

OpenAI 缺乏 DeepMind 那样的谷歌长期资金或庞大员工数量,正在做出一种初创公司式的赌注,即他们知道一个重要的秘密:“Scaling 假设是真的!”因此,像 PPO 这样简单的深度强化学习算法,建立在像 RNN 或 Transformer 这样的大型简单架构之上,可以利用规模的优势涌现出来,并通过元学习获得强大的能力,从而为进一步的算力和 Scaling 获得更多资金,形成良性循环。这就是为什么 OpenAI 不得不修改其公司形式:缺乏像谷歌那样的巨额捐赠或财力极其雄厚的赞助人,它从哪里获得资金来 Scaling(或雇佣年薪数百万的机器学习工程师/研究员)?OpenAI 必须赚取所需的资金,因此,类似于 Mozilla 基金会拥有 Mozilla 公司(以销售 Firefox 搜索引擎位置),或好时孤儿院拥有好时巧克力,或女童子军授权其饼干,OpenAI 从一个纯粹由捐赠资助的非营利组织转变为一个非营利组织拥有营利性子公司/初创公司“OpenAI LP”,后者可以接受投资并从事营利活动。由 OpenAI 控制的 OpenAI LP 然后可以瞄准月球。如果 OpenAI 错误地相信了图表直线之神,那么,他们永远无法直接使用 DeepMind 青睐的方法与 DeepMind 竞争,并且始终只会是一个无足轻重的注脚,所以他们不会后悔。

OA, lacking anything like DM’s long-term funding from Google or its enormous headcount, is making a startup-like bet that they know an important truth which is a secret: “the scaling hypothesis is true!” So, simple DRL algorithms like PPO on top of large simple architectures like RNNs or Transformers can emerge, exploiting the blessings of scale, and meta-learn their way to powerful capabilities, enabling further funding for still more compute & scaling, in a virtuous cycle. This is why OA had to revise its corporate form: lacking any enormous endowment or extremely deep-pocketed patron like Google, where does it get the money to scale (or hire machine learning engineer/researchers who can command salaries in the millions)? OA has to _earn_ the necessary money, so in a move like Mozilla Foundation owning Mozilla Corporation (to sell Firefox search engine placement), or the Hershey orphanage owning Hershey Chocolate or the Girl Scouts licensing their cookies, OpenAI switched from a pure nonprofit funded by donations to a nonprofit which owns a for-profit subsidiary/startup, “OpenAI LP”, which can take investments and engage in for-profit activities. OA LP, while controlled by OA, can then shoot for the moon. And if OA is wrong to trust in the God of Straight Lines On Graphs⁠, well, they never could compete with DM directly using DM’s favored approach, and were always going to be an also-ran footnote, so they have no regret.

虽然所有这些在理论上可以被竞争对手相对容易地复制(永远不要低估所需的调整和特殊配方的数量),如果他们愿意的话(毕竟,所需的算力预算在大科学或其他投资如 AlphaGo、AlphaStar 或 Waymo 面前仍然微不足道),但这些竞争对手缺乏最重要的东西,这是任何金钱或 GPU 都无法治愈的:他们信念的勇气。他们过于保守,在哲学上严重错误,以至于永远无法承认错误并试图超越 OpenAI,直到为时已晚。当美国军方甚至不允许其开发者使用 TensorFlow 或 PyTorch,或者在冠状病毒阴影下的政府项目,我们如何认真谈论任何形式的军事曼哈顿计划?这似乎很荒谬(当然,苦涩的教训/Scaling 假设现在已经积累了足够的先验概率,应该被认真对待,并接受重大研究投资来测试它们能走多远,尤其是考虑到其影响的重要性),但看看每次 OpenAI 发布 Scaling 假设的新例子时受到的反复批评,从 GPT-1 到 Dactyl 到 OA5 到 GPT-2 到 iGPT 到 GPT-3……套用圣奥古斯丁的话,大多数人对苦涩的教训或 Scaling 假设的反应是“赐予我规模和算力——但不是现在”。

While all of this hypothetically can be replicated _relatively_ easily (never underestimate the amount of tweaking and special sauce it takes) by competitors if they wished (the necessary amounts of compute budgets are still trivial in terms of Big Science or other investments like AlphaGo or AlphaStar or Waymo, after all), said competitors lack the very most important thing, which no amount of money or GPUs can ever cure: the courage of their convictions. They are too hidebound and deeply philosophically wrong to ever admit fault and try to overtake OA until it’s too late. How can we talk seriously about any kind of military Manhattan Project when the US military ⁠doesn’t even let its developers use Tensorflow or PyTorch⁠, or about government projects in the shadow of coronavirus? This might seem absurd (surely the Bitter Lesson/scaling hypothesis have now earned enough prior probability to be taken seriously and receive major research investments to test how far they can go, especially given how important the implications are), but look at the repeated criticism of OA _every time_ they release a new example of the scaling hypothesis, from GPT-1 to Dactyl to OA5 to GPT-2 to iGPT to GPT-3… To paraphrase St Augustine, most peoples’ reaction to the Bitter Lesson or scaling hypothesis is “grant me scale & compute—but not yet”.⁠⁠25⁠

一个关键的指标将是,除了“通常的嫌疑对象”(微软 ZeRO-2 团队已达到 1T 规模训练,但还有 Nvidia、Salesforce、Allen、Google DM/GB、Connor/EleutherAI、Facebook FAIR)之外的组织是否开始参与,或者他们是否继续否定 Scaling。至少截至 2020 年 10 月 26 日,152 天后,没有模型接近 GPT-3,实际上,甚至没有模型超过 Turing-NLG 的 17B。

A critical indicator will be whether organizations beyond ‘the usual suspects’ (Microsoft ZeRO-2⁠ team has reached 1t-scale training⁠, but there is also Nvidia, Salesforce, Allen, Google DM/GB, Connor/EleutherAI, Facebook FAIR) start participating or if they continue to dismiss scaling. At least as of 2020-10-26, 152 days later, no model has come near GPT-3, and indeed, no model has even exceeded Turing-NLG’s 17b.⁠⁠26⁠

批评批评者 Critiquing The Critics(https://gwern.net/scaling-hypothesis#critiquing-the-critics "Link to section: § 'Critiquing The Critics'")

事后诸葛亮 权威无问责 寒暄而非预测 官僚铁律:大教堂哥特式

Keeping track Hindsight is 20⁄20 Authority without accountability Phatic, not predictive The iron law of bureaucracy: Cathedral gothic

事后回顾。2020 年的 GPT-3 是回顾过去十年的绝佳节点。回想一下,一个因为对新型“ResNet”感到兴奋而开始攻读博士的人,到现在可能还没毕业——这足以说明 ResNet 本身都是多么近期的事,更不用说 Transformer 了,也体现了进步的速度之快。在 2010 年(16 年前),全世界真正相信深度学习的人可以轻松塞进一个中等大小的会议室(部分原因是其中有 3 人正忙于创立 DeepMind)。2010 年(16 年前)对机器学习感兴趣的人,或许读过一些关于古怪的顽固联结主义者用区区 100 万到 200 万个参数识别手写数字的有趣工作,或者对标准语音识别隐马尔可夫模型进行的一些温和的神经网络改进。在 2010 年(16 年前),谁能预测到未来 10 年深度学习会经历一场寒武纪大爆发,导致机器学习中其他方法的大规模灭绝,模型规模会扩大到 1750 亿个参数,而这些巨大的模型会自发涌现出所有这些能力?

Keeping track. GPT-3 in 2020 makes as good a point as any to take a look back on the past decade. It’s remarkable to reflect that someone who started a PhD because they were excited by these new “ResNets” would still not have finished it by now—that is how recent even resnets are, never mind Transformers, and how rapid the pace of progress is. In 2010 16ya, one could easily fit everyone in the world who genuinely believed in deep learning into a moderate-sized conference room (assisted slightly by the fact that 3 of them were busy founding DeepMind⁠). Someone interested in machine learning in 2010 16ya _might_ have read about some interesting stuff from weirdo diehard connectionists in recognizing hand-written digits using all of 1–2 million parameters, or some modest neural tweaks to standard voice-recognition hidden Markov models. In 2010 16ya, who would have predicted that over the next 10 years, deep learning would undergo a Cambrian explosion causing a mass extinction of alternative approaches throughout machine learning, that models would scale up to 175,000 million parameters, and that these enormous models would just spontaneously develop all these capabilities?

没有人。也就是说,除了少数被 AI 界(更不用说整个世界)视为故意执迷不悟的老派狂热分子的联结主义者,如 Moravec、Schmidhuber、Sutskever、Legg 和 Amodei 之外,没有人。

No one. That is, no one aside from a few diehard connectionists written off as willfully-deluded old-school fanatics by the rest of the AI community (never mind the world), such as Moravec, Schmidhuber, ⁠Sutskever⁠, Legg, & Amodei.

回顾过去最令人震惊的一点是,如果你听对了人,这一切是多么不令人惊讶且容易预测。在 1998 年(28 年前),22 年前,Moravec 指出 AI 研究可能具有欺骗性,硬件限制意味着“智能机器研究在其头 50 年并未稳步前进,而是停滞了 30 年!”,并预测随着摩尔定律的延续,“未来 50 年的进展将比过去 50 年快得多。”Moravec 进一步观察到,快速进展的部分原因是硬件积压:虽然所需算力的超级计算机在联结主义革命开始前很久就已存在,但没有人被允许使用它们,因为它们被用于“更重要”(有声望)的硬 STEM 工作,比如“物理模拟”(即气候模拟和核弹),而“AI 研究必须等待算力变得更便宜。”便宜意味着大约 2229 美元(1998 年)的工作站;足以与人类匹敌的廉价算力将在 2020 年代某个时候到来,而 2010 年代则出现了蜥蜴到老鼠级别的廉价系统。事实上,深度学习革命的起点通常被认为是 2012 年(14 年前)的 AlexNet,由一名研究生使用 2 块 GTX 580 3GB GPU(发布时标价……789 美元(2010 年),系统构建成本约 2285 美元(2012 年))完成。2020 年 GPT-3 问世,如前所述,尽管摩尔定律普遍减速,但除了 2020 年代预计的大幅硬件算力增长外,还有许多理由预期成本会下降。

One of the more shocking things about looking back is realizing how unsurprising and easily predicted all of this was if you listened to the right people. In 1998 28ya, 22 years ago, Moravec noted that AI research could be deceptive, and hardware limits meant that “intelligent machine research did not make steady progress in its first 50 years, it marked time for 30 of them!”, predicting that as Moore’s law continued, “things will go much faster in the next 50 years than they have in the last 50.” Moravec further observed that part of the reason for rapid progress was the hardware overhang: while supercomputers of the necessary power would exist long before the connectionist revolution began, no one would be allowed to use them⁠⁠27⁠, as they would be devoted to ‘more important’ (prestigious) hard STEM work, like “physics simulations” (ie. climate simulations & nuclear bombs)⁠⁠28⁠, and “AI research must wait for the power to become more affordable.” Affordable meaning a workstation roughly ~$2,229$1k 1998; sufficiently cheap compute to rival a human would arrive sometime in the 2020s, with the 2010s seeing affordable systems in the lizard–mouse range. As it happens, the start of the DL revolution is typically dated to AlexNet⁠ in 2012 14ya, by a grad student⁠⁠29⁠ using 2 GTX 580 3GB GPUs (launch list price of… $789$500 2010, for a system build cost of perhaps $2,285$1,500 2012). 2020 saw GPT-3 arrive, and as discussed before, there are many reasons to expect the cost to fall, in addition to the large hardware compute gains that are being forecast for the 2020s despite the general deceleration of Moore’s law.⁠⁠30⁠

过去 10 年加速的进展应该将任何人从教条式的沉睡中唤醒,让他们坐直身子。事实证明,这是 Hans Moravec 的世界,而我们其他人只是生活在一个愚人的天堂。Moravec 的预测还有 28 年……

The accelerating pace of the last 10 years should wake anyone from their dogmatic slumber and make them sit upright. It turned out, it’s Hans Moravec’s world, and the rest of us were just living in a fool’s paradise. And there are 28 years left in Moravec’s forecast…

许多人不仅不抵制,反而沉溺于一种诱惑,即屈服于职业偏见,将任何模型贬低为“不过是”这个或那个(“不过是数十亿条 IF 语句”、“不过是一堆乘法”、“不过是数百万个记忆的网页”),只见树木不见森林,正如 Moravec 对国际象棋引擎的评论:

The temptation, that many do not resist so much as revel in, is to give in to a _déformation professionnelle_ and dismiss any model as “just” this or that(“just billions of IF statements” or “just a bunch of multiplications” or “just millions of memorized web pages”), missing the forest for the trees, as Moravec commented of chess engines:

但当然,如果我们成功实现了 AI,或者一般意义上的还原论,那必然是通过将 Y 还原为“不过是 X”。证明某个需要智能的任务可以通过一个定义明确的、没有“智能”的算法来解决,这正是成功必须呈现的样子!(否则,问题就被彻底回避了,只是被推到了别处;计算机芯片是由晶体管构成的,而不是特别小的侏儒。)

But of course, if we ever succeed in AI, or in reductionism in general, it _must be by reducing Y to ‘just X’_. Showing that some task requiring intelligence can be solved by a well-defined algorithm with no ‘intelligence’ is precisely what success must look like! (Otherwise, the question has been thoroughly begged & the problem has only been pushed elsewhere; computer chips are made of transistors, not especially tiny homunculi.)

事后诸葛亮。即使在 2015 年(11 年前),所有专家都向我们保证,AGI 的缩放假说似乎非常可疑:毕竟你需要有东西可以缩放,而且很容易看到现有系统的缺陷,想象它们永远不会消失,进展随时会趋于饱和。就像基因组学革命,少数有远见的先知推断 GWAS 所需的样本量将呈指数增长,并很快产生强大的多基因评分,而清醒的专家则对“缺失的遗传力”和生物学奇迹般的复杂性感到焦虑,嘲笑如此大的样本量要求证明了 GWAS 是一个失败的范式,未来先是缓慢到来,然后迅速到来。然而,我们现在就在这里:向狂热分子致敬,批评者应感到羞耻和羞辱!如果能回到 10 年前,甚至 5 年前,看着每个 AI 研究人员读到这篇论文时脑袋爆炸……不幸的是,现在似乎没有多少脑袋爆炸,因为人类事后诸葛亮和找借口的能力是无限的(“我微调一下也能得到那么多,反正我早就预测到了,多无聊”),而且不幸的是,“没有火警警报”来预警 AGI。(如果你仍然确信未来几十年 AGI 的概率接近零,为什么?你是否——以书面形式——预测过像 GPT-3 这样的能力?这就是你预期 AI 失败在之前几十年会呈现的样子吗?什么样的具体任务、什么样的具体数字才能说服你?如果这些原始的、昆虫大脑大小的深度学习系统并非通往成功的道路,那么世界会与现在有什么不同?)

Hindsight is 20⁄20. Even in 2015 11ya, ⁠all the experts⁠ assured us that AGI the scaling hypothesis seemed highly dubious: you needed something to scale, after all, and it was all too easy to look at flaws in existing systems and imagine that they would never go away and progress would sigmoid any month now, soon. Like the genomics revolution where a few far-sighted seers extrapolated that the necessary _n_ for GWASes would increase exponentially & deliver powerful PGSes soon, while sober experts wrung their hands over “missing heritability” & the miraculous complexity of biology & scoff about how such _n_ requirements proved GWAS was a failed paradigm, the future arrived at first slowly and then quickly. Yet, here we are: all honor to the fanatics, shame and humiliation to the critics!⁠⁠31⁠ If only one could go back 10 years, or even 5, to watch every AI researchers’ head explode reading this paper… Unfortunately, few heads appear to be exploding now, because human capacity for hindsight & excuses is boundless (“I can get that much with finetuning, anyway I predicted it all along, how boring”) and, unfortunately, “there is no fire alarm”⁠ for AGI. (If you are still _certain_ that there is near-zero probability of AGI in the next few decades, why? Did you predict—in writing—capabilities like GPT-3? Is this how you expect AI failure to look in the decades beforehand? What specific task, what specific number, would convince you otherwise? How would the world look different than it does now if these crude prototype insect-brain-sized DL systems were not on a path to success?)

权威无问责。我们应该如何看待专家?失败的预测是由杰出、可敬、严肃的人做出的。他们以深思熟虑的口吻谈论为什么 AI 炒作过度,可能引发“AI 寒冬”,以及流行方法的根本缺陷,为什么蛮力行不通。这些言论在 2014 年(12 年前)、2015 年(11 年前)、2016 年……屡见不鲜。而他们都错了。据我所知,很少有人承认错误或进行反思。这是一个令人费解的失败,我以前也反思过。

Authority without accountability. What should we think about the experts? Projections of failure were made by eminent, respectable, serious people. They spoke in considered tones of why AI hype was excessive and might trigger an “AI winter”, and the fundamental flaws of fashionable approaches and why brute force could not work. These statements were made routinely in 2014 12ya, 2015 11ya, 2016… And they were wrong. I am aware of few issuing a _mea culpa_ or reflecting on it.⁠⁠32⁠ It is a puzzling failure, and I’ve ⁠reflected on it before⁠.

寒暄而非预测。然而,有一种特定的语调,所有自以为是的人都在使用,无论对错听起来都一样;这种语调与今年 1 月到 3 月的许多言论相同;我们也可以在 1940 年(86 年前)一篇《科学美国人》文章中找到这种语调,该文章权威地题为“别担心——这不可能发生”,建议读者不要再为此担心,“去睡觉吧”。(“它”指的是原子弹,某些科学家已停止谈论它,引发了公众担忧;不仅可能发生,英国原子弹项目已经启动,5 年后它确实发生了。)

Phatic, not predictive. There is, however, a certain tone of voice the bien pensant all speak in, whose sound is the same whether right or wrong; a tone shared with many statements in January to March of this year; a tone we can also find in a 1940 86ya _Scientific American_ article authoritatively titled, “Don’t Worry—It Can’t Happen”⁠, which advised the reader to not be concerned about it any longer “and get sleep”. (‘It’ was the atomic bomb, about which certain scientists had stopped talking, raising public concerns; not only could it happen, the British bomb project had already begun, and 5 years later it did happen.)

官僚铁律:大教堂哥特式。这种语调就是权威的语调。

The iron law of bureaucracy: Cathedral gothic. This tone of voice is the voice of authority.

权威的语调坚持要冷静,人们不要“恐慌”(罪中之首)。

The voice of authority insists on calm, and people not “panicking” (the chief of sins).

权威的语调向你保证它不会发生(因为它不可能发生)。

The voice of authority assures you that it won’t happen (because it can’t happen).

权威的语调提出关于现状将占上风的简单论点,只考虑新想法可能如何失败(而不考虑所有可能的选项)。

The voice utters simple arguments about why the status quo will prevail, and considers only how the wild new idea could fail (and not all the possible options).

权威的语调不涉及不确定性;事情要么发生要么不发生,既然它不会发生,就没有必要采取任何预防措施(你不应该担心,因为它不可能发生)。

The voice is not, and does not deal in, uncertainty; things will either happen or they will not, and since it will not happen, there is no need to take any precautions (and you should not worry because it can’t happen).

权威的语调不相信在图表上画线(那是纯粹的命理学)。

The voice does not believe in drawing lines on graphs (it is rank numerology).

权威的语调不做任何数值预测(这些预测可能被证伪)。

The voice does not issue any numerical predictions (which could be falsified).

权威的语调不会分享其源代码(原因复杂,无法向外行解释)。

The voice will not share its source code (for complicated reasons which cannot be explained to the laity).

权威的语调反对不道德的事情,比如对志愿者进行随机实验(但会忽略侮辱)。

The voice is opposed to unethical things like randomized experiments on volunteers (but will overlook the insult).

权威的语调没有未来的模型(因为模型意味着它尚未知道未来)。

The voice does not have a model of the future (because a model implies it does not already know the future).

权威的语调关心其公众形象(以及由其他权威语调者对其进行的刻薄八卦)。

The voice is concerned about its public image (and unkind gossip about it by other speakers of the voice).

权威的语调总是清醒、可敬、有资质的(权威语调很乐意为你国家的杂志和/或报纸撰写专栏)。

The voice is always sober, respectable, and credentialed (the voice would be pleased to write an op-ed for your national magazine and/or newspaper).

权威的语调说话,而不是被说话(你不能问权威语调什么客观事实会改变它的想法)。

The voice speaks, and is not spoken to (you cannot ask the voice what objective fact would change its mind).

权威的语调从不改变主意(直到它改变)。

The voice never changes its mind (until it does).

权威的语调从不对世界上的事件感到惊讶(只感到失望)。

The voice is never surprised by events in the world (only disappointed).

权威的语调建议你回去睡觉(现在)。

The voice advises you to go back to sleep (right now).

当有人谈论未来的可能性时,他们说话的语调是什么?

When someone speaks about future possibilities, what is the tone of their voice?

万物源于比特 It From Byte(https://gwern.net/scaling-hypothesis#it-from-byte "Link to section: § 'It From Byte'")

我之前曾论证过,GPT-3 显然展现出智能体性,因为它从人类生成的文本数据中进行离线模仿学习(具体来说是行为克隆),从而学习到许多智能体(真实或虚构)的生成模型。这些生成模型提供了智能体式能力,因为它们可用于提示模型进行“角色扮演”——规划并采取行动,将环境引导至状态空间中的小目标区域;这并非仅仅是假设性的,也不局限于其内部模拟环境中行动与结果的文本记录,而是当赋予效应器时,例如在 SayCan 案例中,语言模型实际上会在现实世界中执行此类操作。

I have previously argued that GPT-3 clearly shows agency because it is doing offline imitation learning (behavioral cloning, specifically) from the human-generated text data, and so it learns generative models of many agents, real or fictional. These generative models offer agentic capabilities, because they can be used to prompt the model to ⁠‘roleplay’⁠—plan & take action which will steer environments into small goal regions of state-space; and this is not merely hypothetical, or confined to text transcripts of actions & results in its internal simulated environments, but given effectors, like in the case of SayCan⁠, a language model will in fact do such things in the real world.

这类系统可能从未“体验过真实世界”,也未曾被刻意训练过恶意智能体的精确动作序列,但这并不意味着它们无法泛化或模仿。一个足够精确的智能体模拟本身就是一个智能体。(我们可以为 GPT-3 设置一个提示来模仿阿道夫·希特勒,询问他如何重新掌权并继续灭绝犹太人,然后得到一个半连贯的高层计划;这很不幸,而且模拟对象甚至不必是真实人物——虚构的邪恶角色同样能轻易策划邪恶之事,因为想象他们“会”想做什么可怕的事情并不困难。)这似乎与公认的强化学习实例(如行为学习或离线强化学习)并无太大区别:如果你从智能体的数据(无论是人类还是 DRL 智能体的日志数据)进行训练,那么问题就是“你如何能不从这些例子中学会如何行动并具备追求目标的能力?” 除非你是一个愚蠢的模型,规模太小或数据太少,否则你大概无法避免。

That such systems may never have ‘experienced the real world’ or been trained deliberately on exact action sequences of malicious agents doesn’t mean that they cannot generalize or imitate. A sufficiently accurate simulation of an agent just _is_ an agent. (One can set up a prompt for GPT-3 to imitate Adolf Hitler and ask him how to regain power & resume exterminating the Jews and get back a semi-coherent high-level plan; this is unfortunate, and the simulacra need not even be of a real person—evil fictional characters plan evil things just as easily, because it’s not hard to imagine what horrible things they _would_ want to do.) This doesn’t seem all that different from accepted instances of reinforcement learning, like behavior learning or offline reinforcement learning: if you train on data from agents, whether humans or logged data from DRL agents, then the question is “how would you _not_ learn from all these examples how to act & be capable of pursuing goals?” Presumably you would not only if you were a stupid model, too small or given too little data.

如果这些都不是“智能体”,我不知道什么才“真正”是智能体;或者至少,如果批评者坚持某种排除这些系统的“智能体”定义,我认为我们或许应该完全放弃“智能体”这个词——因为如果给 SayCan 机器人一个指令“去拿一罐可乐并带给我”,它利用图像输入构建逐步计划来寻找、获取并带回可乐罐,并且在真实机器人上经常成功做到这一点,这都不算“智能体”,那么我们需要一个词来描述这种非智能体系统,以便讨论它们的非智能体危险。(如果我们因为缺乏附属肢体而将其定义为子智能体,从而将所有模型定义为无害的非智能体,鉴于人们在第一时间就将模型连接到人类、API、搜索引擎或机器人时所表现出的极度粗心和漫不经心,这是一种不可接受的混淆——OpenAI 的 GPT-3 API 在 2020 年 7 月刚推出,人们就炫耀利用其基本的 HTML/CSS/JS 能力来驱动网页浏览器,而像 LaMDA 或 Adept 这样的大型语言模型开发者则表现出一种不恰当的急切,让模型查询任意 URL,甚至在其论文中都没有说明这是实时的。AI 盒子还没发明出来,每个人就决定让他们的 AI 走出盒子以变得稍微更有用,这毫不意外——毕竟,工具 AI“想要”成为智能体 AI。)

If these are not ‘agents’, I don’t know what “really” is an agent; or at least if critics insist on some sort of definition of ‘agent’ which excludes these, I think perhaps we should then abandon the word ‘agent’ entirely—because if giving a SayCan robot an instruction to ‘fetch a can of Coke and bring it to me’, with it using image inputs to construct step-by-step plans to find, possess, and return with the can, and successfully doing so often in real life on a real robot, does not count as an ‘agent’, then we need a word for such non-agent systems, so we can discuss their non-agency dangers. (If we define them as sub-agents because of lack of appendages and thus define all models as harmless non-agents, this is an unacceptable equivocation given the extreme carelessness and insouciance people display in hooking up their models the first chance they get to humans, APIs, search engines, or robots—hardly had the OpenAI GPT-3 API been launched in July 2020 than people were showing off using its basic HTML/CSS/JS abilities to drive web browsers, and large LM model developers like LaMDA or Adept display an unseemly eagerness to let it query arbitrary URLs without their paper even bothering to specify it was live. The AI box hadn’t even been invented before everyone decided to let their AI out of the box to be slightly more useful, as should come as no surprise—after all, tool AIs _want_ to be agent AIs⁠.)

但人们可能想知道这种逻辑能走多远:我们的工具 AI 中涌现出智能体 AI,仅仅是因为我们在大量智能体生成的数据上训练了它们吗?如果我们抛弃人类文本语料库(其中充满了关于人类规划、行动和实现目标的文本),以及充满智能体活动的视频数据集,并且也删除图像数据集(因为它们只是视频的快照,描绘了智能体、行动以及充满智能体痕迹的环境),那么我们是否会得到一个相对安全的“工具 AI”模型,不再有隐藏的智能体性?

But one might wonder how far this logic goes: do we have agent AIs emerging from our tool AIs _only_ because we trained them on so much agent-generated data? If we scrapped human text corpuses, full of text about humans planning and taking actions and obtaining goals, or video datasets stuffed full of agents doing stuff, and if we deleted image datasets as well because they are just snapshots of videos and depict agents & actions & environments full of traces of agency, would we then have a model which is now just a (relatively) safe ‘tool AI’, with no agency lurking?

我仍然认为存在这种可能性,而且可能性或许并不小:智能体性不是一个离散的事物,而是一个连续体,它是一种收敛的工具性驱动力/涌现能力,因为即使对于理解“非智能体”事物也是有用的。

I would still say that there’s a possibility, and maybe not even that small one: agency is not a discrete thing, but a continuum, which is a convergent instrumental drive / emergent capability because it is useful even for understanding “non-agentic” things.

* GPT-3 创意小说(完整上下文):

* GPT-3 Creative Fiction⁠ (⁠full context⁠):

万物皆原子与虚空 All Is Atoms & Void(https://gwern.net/scaling-hypothesis#all-is-atoms-void "Link to section: § 'All Is Atoms & Void'")

首先,在“智能体式”数据和仅仅是“自然”的数据之间,不可能存在任何原则性的、严格且必要的区分。这是因为在现实中也不存在这样的区分:所有“智能体”都是由非智能体的原子等基本单元构成的。不存在智能体粒子,也没有松果体提供通往“_真正决策™_”的通道。一个智能体的人类与太空中旋转的一团尘埃、一块岩石或一台计算机由相同的基本物质构成。因此,仅从原子(可能很多)的模拟出发,除了最原始的物理方程和原子与虚空之外别无他物,最终重现宇宙的历史并观察到诸如生命起源和人类等事物,这一定是可能的。因此,人们将非智能体数据(物理方程)转化为智能体数据。

First, there cannot be any principled, hard-and-fast, necessary distinction between data which is ‘agentic’ and data which is merely ‘natural’. This is because there is no such distinction in reality either: all ‘agency’ is constructed of non-agentic bits like atoms. There is no agency-particle, no pineal gland granting access to ‘_Genuine_ Decision-Making™’. An agentic human is made out of the same fundamental things as a clump of dust swirling in space, or rock, or a computer. It must be the case that one could, starting only from simulations of (possibly a lot of) atoms, nothing but the most raw physics equations and atoms & the void, and eventually recapitulate the history of the universe and observe things like the origin of life and humans. Thus, one turns non-agentic data (physics equations) into agentic data.

好吧,但除非有超级计算机,否则这不太可能发生。如果我们考虑现实水平的算力,比如当代的神经网络,训练数据少于宇宙中的一切,且表面上无害,例如河流向下流动的水文学(如防洪),或太阳系的轨迹,那么肯定不会有任何智能体演化出来——无论对冥王星的混沌动力学进行多少建模,都不会对关于冥王星是否是行星的天文学内讧的动力学建模有任何帮助,对吧?

OK, but barring a hypercomputer, that is unlikely to happen. If we consider realistic levels of compute, like contemporary NNs, trained on less-than-everything-in-the-universe & apparently harmless data like, say, the hydrology of rivers flowing downhill (eg. for flood prevention), or the trajectory of the solar system, surely none of that agency will evolve—no amount of modeling the chaotic dynamics of Pluto will give you any help in modeling the dynamics of astronomy infighting about whether Pluto is a planet, right?

意向性解释立场 Intentional Interpretive Stance(https://gwern.net/scaling-hypothesis#intentional-interpretive-stance "Link to section: § 'Intentional Interpretive Stance'")

在此我再次持不同意见,并援引丹尼尔·丹尼特的意向性立场。事实上,人类确实会将此类自然系统建模为智能体。我们发现,对于许多自然系统而言,这种目的论解释对于直觉和捷径推理是不可或缺的。

Here again I differ, and invoke Daniel Dennett’s⁠intentional stance⁠. Humans do, in fact, model natural systems like these as agents. We find such teleological explanations indispensable for intuition and shortcut reasoning across many natural systems.

变分解释 Variational Interpretations(https://gwern.net/scaling-hypothesis#variational-interpretations "Link to section: § 'Variational Interpretations'")

Janus 评论道,针对他们强调的、我可能称之为 GPT-3 的‘以世界模型为中心’的直觉,与我‘以智能体为中心’的观点形成对比:

⁠Janus comments⁠, apropos of their emphasis on what I might call a ‘world-modeling-centric’ intuition for GPT-3 vs my ‘agent-centric’ view that:

我接受这种描述:这实际上很自然,并非地心说世界模型上的复杂本轮,而是日心说——强大、有用且更简单。它(同样像日心说³³)可能违反直觉,这令人遗憾,但其优点已被证明。

I embrace that description: it is in fact natural and not an elaborate epicycle on a geocentric model of the world, but rather, heliocentrism—powerful, and useful, and simpler. That it (also like heliocentrism⁠⁠33⁠) may feel counterintuitive is unfortunate, but its virtues are proven.

如果我们采取意向立场,陷入感伤谬误,说河神想要与海洋重聚(因此我们必须献祭以免堤坝决口),那就错了;但当我们说河流试图找到最小化其重力或自由能的最优路径时,我们是对的。这既真实、预测有用,又在数学上等价于另一种表述方式,即通过‘前向’过程逐步、逐原子计算,并得到相同答案——但通常更容易求解。(特德·姜的《你一生的故事》试图通过小说传达这一视角。)这种捷径是我们普遍使用的技巧,适用于从河流下坡到行星轨道,再到光子穿过水时最小化传播时间的路径,乃至进化动力学:不是试图逐步理解,而是通过变分原理将系统视为整体,‘想要’最小化(或最大化)某个简单的全局量(奖励),并选择实现这一目标的动作序列。(‘河流想要最小化其高度,因此无需模拟到单个水流,我只需查看地图,就能看出它应该“选择”向左,然后向右,再蜿蜒穿过这片略微倾斜的平坦区域。啊,看来我是对的。’)然后,将这个模块化技巧应用于具体系统和量,像智能体一样思考……³⁴

We err if an intentional stance leads us engage in the pathetic fallacy and say that the river-spirit wants to reunite with the ocean (and we must offer sacrifices lest the dikes breach), but we are correct when we say that the river tries to find the optimal path which minimizes its gravitational or free energy⁠. It is both true, predictively useful, and mathematically equivalent to the other way of formulating it, in terms of ‘forward’ processes computing step by step, atom by atom, and at the getting the same answer—but typically much easier to solve. (Ted Chiang’s “Story Of Your Life”⁠ tries to convey this perspective via fiction.) And this shortcut is a trick we can use universally, for everything from a river flowing downhill to the orbit of a planet to the path of photon through water minimizing travel time⁠ to evolutionary dynamics: instead of trying to understand it step by step, treat the system as a whole via the variational principle⁠ as ‘wanting’ to minimize (or maximize) some simple global quantity (a reward), and picking the sequence of actions that does so. (“The river _wants_ to minimize its height, so without simulating it down to the individual water currents, I can look at the map and see that it should ‘choose’ to go left, then right, and then meander over this flat slightly-sloping part. Ah, looks like I was right.”) Then, into this modular trick, just plug in the system and quantity in question, and think like an agent…⁠⁠34⁠

哦哦。‘预测有用’、‘捷径’、‘容易得多’、‘普遍适用’——这些都是神经网络喜爱的特性。对它来说很自然。当存在一个 AI 安全研究人员讨厌的奇怪技巧,比如采用目的论和变分推理时,它为什么要用一套单独、计算昂贵的技巧来解决每个异质问题呢?

Uh oh. ‘Predictively useful’, ‘shortcut’, ‘much easier’, ‘universally’—all properties a neural net loves. All natural to it. Why would it try to solve each heterogeneous problem with a separate, computationally-expensive, bag of tricks, when there’s one weird trick AI safety researchers hate, like adopting teleological and variational reasoning?

诱导涌现代价高昂 Inducing Emergence Is Expensive(https://gwern.net/scaling-hypothesis#inducing-emergence-is-expensive "Link to section: § 'Inducing Emergence Is Expensive'")

当然,这种框架可能比直接解决问题更昂贵。变分方法虽然强大但反直觉,而且模型通常可以采用许多更简单的近似或记忆方法。对于像冥王星轨道建模这样的单一问题,不太可能学到任何变分方法。当只有一个系统和单一最小化量时,它们可以直接被假设,何必学习变分方法?这类似于预训练诱导的其他模型能力:诸如归纳头、元学习、计数或推理等能力需要证明其价值,并且从一开始并不优于替代方案。它们需要足够丰富的模型来可行地计算,足够的数据来迫使它们摆脱更简单的解决方案(这些方案会在少数罕见数据点上失败),以及足够的训练(以遍历所有可能性,收敛到更好的能力)。

Of course, this frame can be more expensive than solving a problem directly. Variational approaches are powerful but counterintuitive, and there are often many simpler approximations or memorization that a model can do. For a _single_ problem like modeling the orbit of Pluto, it is unlikely that any variational approach would be learned. Why would it, when there is only 1 system and 1 quantity minimized, so they can just be assumed? This is similar to other ⁠model capabilities induced by pretraining⁠: things like induction heads or meta-learning or counting or reasoning need to pay their way, and are not superior to alternatives right from the start. They need rich enough models to compute them feasibly, enough data to force them out of easier solutions (which will fail on a few rare datapoints), and enough training (to work through all the possibilities to converge on the better capabilities).

什么能诱导智能体涌现? What Can Induce Agency Emergence?(https://gwern.net/scaling-hypothesis#what-can-induce-agency-emergence "Link to section: § 'What Can Induce Agency Emergence?'")

不幸的是,这是一个经验问题。需要多少个数据集?每个数据集有多大?它们需要有多多样化?甚至什么是“数据集”,因为我们总是可以合并或拆分?我们很难预测 GPT-3 中何时会出现某种能力,所以我们绝对无法先验地说“冥王星是安全的建模对象,但加入几千个系外行星系统就会开始引发一个定义系统/插入奖励/最大化模块并带回智能体”。

Unfortunately, this is an empirical matter. How many datasets? How big is each dataset? How diverse do they have to be? What even is a ‘dataset’, since we can always lump or split it? We struggle to predict when a capability will develop in GPT-3, so we definitely can’t say a priori that “Pluto is safe to model, but then tossing in a few thousand exoplanet solar systems would begin to elicit a define-system/plug-in-reward/maximize module and bring back agency”.

也很难说清楚哪些数学或物理系统表现出正确的最大化行为,从而可以泛化到意向立场。超抽象且简单的元胞自动机——康威生命游戏(GoL)——能否诱导出意向立场?

It would also be hard to say at all what mathematical or physical systems exhibit the right kinds of maximizing behavior which can be generalized to an intentional stance. Does the ultra-abstract & simple cellular automaton⁠Conway’s Game of Life⁠ (GoL) induce an intentional stance?

它没有智能体,没有生物学,没有通常意义上的进化——但它确实有许多小模式,这些模式可以被有用地进行分块,以帮助理解特定的 GoL。当然,人类将 GoL 视为一堆像“滑翔机”这样的小实体,但一个给定随机初始化棋盘的神经网络也可能看到同样的东西,因为大多数 GoL 模式会消亡或达到像滑翔机或静物模式这样的不动点:将一个大 GoL 棋盘分块成几个在“空空间”中游荡、偶尔被“静物”打断的“滑翔机”只是更简单而已。

It has no agents, no biology, no evolution in the usual sense—but it does have many small patterns which can be usefully chunked⁠) to help understand a specific GoL. Humans, of course, look at a GoL as a bunch of small entities like ‘gliders’, but a NN given randomly-initialized boards may also see the same thing, because most GoL patterns will die out or reach fixed-points like gliders⁠) or still-life⁠) patterns: it is simply simpler to chunk a large GoL board into a few ‘gliders’ wandering through ‘empty space’ interrupted by the occasional ‘still life’.

一旦你谈论滑翔机在游荡,除非它们撞上会杀死它们的静物块,你就已经很大程度上接近了意向立场——不是将滑翔机建模为将局部邻域的这条那条规则应用于一百万个同等重要的细胞的不可阻挡的结果,而是将其视为一个特定的感兴趣实体,相对于隐含且被忽略的死亡细胞背景,它将四处移动并射向无穷远或遭遇末日。

And once you are talking about gliders wandering around unless they run into a still-life block which kills them, you are much of the way to an intentional stance—not modeling a glider as the inexorable outcome of applying this and that rule about the local-neighborhood to a million cells of equal importance, but as a specific entity of interest against an implicit & ignored background of dead cells, which will travel around and shoot off to infinity or meet its doom.

所以,我不太愿意打赌 GoL 无法诱导任何迁移。

So, I wouldn’t want to bet too much on GoL being unable to induce any transfer.

我们能走得更广吗?比如,不是自然物理系统,也不是人类感兴趣的特定抽象(GoL 在元胞自动机中特别有趣,我们忽略了定义了一个 CA 但什么有趣的事都不做的大量 CA 规则空间),而是所有图灵机,假设我们采用随机规则,并带有一个长度偏置的随机程序样本,我们将其交叉并视可用磁带为序列预测问题。毕竟,没有比这更通用的可计算设置了。

Can we go even broader? How about, not natural physics systems, nor specific abstractions of interest to humans (GoL is especially interesting among cellular automatons, and we ignore the large space of CA rules which define a CA but which does nothing interesting), but all Turing machines, let’s say random rules with some length-biased sample of random programs which we dovetail & treat available tapes as a sequence prediction problem. There is no more general computable setting, after all.

在随机图灵机上训练是否会带来智能体的可能性?

Would training on a random Turing machine risk the possibility of agency?

也许不会。对于单个 TM,这可能会培养一些能力,比如指令遵循(出于同样的原因,在源代码上预训练,特别是带有状态日志的源代码,是许多任务的强大先验),但它似乎没有任何会诱导智能体的特征。随机 TM 程序不会试图最小化或最大化任何东西;它们只是运行。它们不会试图最大化运行时间长度(终止或不终止),或者尽可能少或尽可能多地在磁带上写入,或者实现特定模式。模型只会学习 TM 规则,并尝试在其有限的前馈神经网络资源下尽可能好地近似它;最终,如果它能迭代或循环地工作,它将学会规则并完美泛化,不再有进一步的学习。通过是否停机来分类 TM 程序没有帮助:是的,忙碌的海狸“想要”最大化某些东西,但这只是定义上作为最长的终止程序,还有更多 TM 程序“乐于”很快停机。所以预测停机状态可能会学到一些东西,但也没有任何表面上看起来像智能体的东西。

Maybe not. For a single TM, this might foster some capabilities like instruction-following (for the same reason that pretraining on source code, especially source code augmented with state logs, is a powerful prior for many tasks), but it does not seem to have any of the traits that would induce agency. There is nothing that random TM programs try to minimize or maximize; they simply run. They don’t try to maximize run time length (terminating or non-terminating), or write as few or as many places on the tape as possible, or to achieve particular patterns. A model would simply learn the TM rules and attempt to approximate it as best as it can given its own limited feedforward neural net resources; eventually, if it can work iteratively or recurrently, it would learn the rules and generalize perfectly, and no further learning occurs. Classifying TM programs by whether they halt doesn’t help: yes, the Busy Beaver ‘wants’ to maximize something, but that’s just by definition as the longest terminating program, there are many more TM programs which are ‘happy’ to halt very quickly. So predicting halting status may learn things, but also still nothing that prima facie looks like agency.

这可能是因为只有一个 TM,类似于只在冥王星上训练。也许正确的设置是在许多 TM 规则(以及每个规则内的程序)上进行训练。这正是研究人员更感兴趣的,因为很少有 TM 具有任何内在兴趣,我们也不知道那个唯一的真图灵机™;我们更希望有一个神经网络在学习学习 TM,或者元学习,而在来自一个分布的许多环境上训练神经网络是诱导元学习的最简单方法。那么,如果我们训练一个模型对随机 TM 加随机程序进行序列预测,而不重复使用呢?如果单个随机图灵机是无害的,那么所有图灵机呢?

This might be due to there being only a single TM, making it analogous to training only on Pluto. Perhaps the right setting would be training over _many_ TM rules (and programs within each one). This is what a researcher would be more interested in, since few TMs are of any intrinsic interest, nor do we know the One True Turing Machine™; we’d rather have a neural network which is learning to learn TMs, or meta-learning, and training a NN over many environments drawn from a distribution is the easiest way to induce meta-learning. So what if we trained a model to do sequence prediction of a random TM + random program, without reuse? If single random Turing machines are harmless, how about all of them?

嗯,好吧……值得注意的是艾伦·图灵是如何引入图灵机形式化的:作为一个通用设置,其中一个人阅读并执行关于如何在纸带上做标记的规则集。所以即使在计算机作为仅仅执行编程指令的工具的原始表述中,我们也有一个侏儒在中心!这个侏儒可以对磁带做任何事情(并且给定不同的指令,他会做),但他想要准确地遵循当前的指令集,直到完成。在每次从 TM+程序分布中抽取时,他遵循不同的指令集,现在模型试图推断他想要什么,以便尽可能快地通过重新计算来准确预测序列。

Hm, well… It’s worth noting how Alan Turing introduced the Turing machine formalism: as a general setting in which a _man_ read and executed sets of rules about how to mark up a paper tape. So even in the original formulation of computers as tools which merely do what they are programmed to do, we have a homunculus at the center! This homunculus could do (and given different instructions, would) anything to the tape, but he wants to follow the current set of instructions accurately, until he’s done. In each draw from the TM+program distribution, he is following a different set of instructions, and now the model is attempting to infer what he wants, to as quickly as possible begin predicting the sequence accurately by recomputing it.

这提供了我们的模块化,以及一个特定的计算执行,以及强烈的优化压力来快速“读取”历史并推断新规则必须是什么。这可能没有干净的最大化奖励的解释,但它确实听起来很像任何人对任何类型的智能体所做的事情:推断奖励函数的逆强化学习问题可能任意困难,在我们成功之前,我们反而推断局部规则和模式,这些规则和模式针对特定的结果(状态空间区域)。你可能不知道为什么你的邻居会做他做的那些奇怪的事情,但你可以推断他会做,而不是另一个智能体,甚至不是他邪恶的双胞胎兄弟。推断 TM 规则是最简单、最原始的“心智理论”吗?也许吧。在这种情况下,无处可逃智能体的可能性。

This provides our modularity, and a particular computation executed, and strong optimization pressure to rapidly ‘read’ the history and infer what the new rules must be. That may not have a clean reward-maximizing interpretation, but it _does_ sound a lot like what anyone does with an agent of any kind: the inverse reinforcement learning problem of inferring the reward function can be arbitrarily hard, and until we succeed at that, we instead infer local rules & patterns, which target particular outcomes (regions of state-space). You may not know why your neighbor does that weird thing he does, but you can infer that he will do it, and not another agent, not even his evil identical twin. Is inferring TM rules the simplest & most rudimentary possible ‘theory of mind’? Maybe. In which case, there is no escape from the possibility of agency anywhere.

环境智能体 Ambient Agency(https://gwern.net/scaling-hypothesis#ambient-agency "Link to section: § 'Ambient Agency'")

智能体可能类似于图灵完备性:即使在缺乏选择或优化的环境中,它也是一种过于有用且过于趋同的能力,以至于无法保证其不存在。系统越广泛、越强大,下一个特征或下一批数据就越可能将其推过临界点,并且设计一个没有该方面的系统就变得更加困难。

Agency may be like Turing-completeness⁠: even in settings free of selection or optimization, it is a capability too useful and too convergent to guarantee its absence. The broader and more powerful a system is, the more the next feature or next piece of data may push it over the edge, and it becomes harder to engineer a system _without_ that aspect.

智能体可以从智能体生成的数据中学习,这些数据具有极强的选择性。或者,如果你仔细移除所有这些数据,它可能来自非人类数据的选择。或者,它可能隐含在复制系统的动力学中。或者,它可能是无数物理系统之一,这些系统具有这样的解释,这些解释在计算上更高效,因此任何优化以平衡可实现算力与准确性的神经网络都会被推向这些解释。或者,它可能是具有宏观统计的系统的良好简化,其中详细的微观状态贡献甚微。或者,它可能仅仅源于在图灵机上的规则归纳的元学习,因为智能体可能遵循复杂的学习策略集,但激励的奖励函数是一个欠定的黑箱。

Agency can be learned from data generated by agents, who generate extremely selective data. Or if you carefully remove all that, it may come from the selection of non-human data. Or it may be implicit in the dynamics of a replicator system. Or it may be one of the countless physical systems which have such interpretations which are computationally more efficient and thus any NN which is optimized to balance realizable compute with accuracy will be pushed to such interpretations. Or it may be a good simplification of systems with macro-statistics where the detailed micro-state adds little. Or it may stem from simply meta-learning of rule induction on TMs, because agents may follow complex sets of policies which are learnable but the motivating reward-function is an under-determined blackbox.

或者……就像压制图灵完备性一样,一旦沉船上的一个洞被堵住,你就会注意到另一个漏洞出现。你无法压制一个好想法。你所能做的就是制造一个复杂的系统,就你所知,它不表现出智能体性;不幸的是,就像图灵完备性(或安全漏洞)一样,没有明显的智能体性并不意味着它不存在。模型不会告诉你,它只是在忙于降低损失。(“采样可以显示知识的存在,但不能显示其不存在。”)

Or… like squashing Turing-completeness, as soon as one hole in the sinking ship is patched, you notice another leak spring up. You can’t keep a good idea down. All you can do is make a complex system that doesn’t display agency as far as _you_ can tell; unfortunately, much like Turing-completeness (or security vulnerabilities), that there is no overt agency doesn’t mean it is not there. The model won’t tell you, it is just getting on with the job of lowering its loss. (“Sampling can show the presence of knowledge, but not the absence.”)

对此我没有任何解决方案,只能再次建议放弃那个诱人、方便但错误的想法,即工具型 AI(无论其品牌是‘工具型 AI’、‘物理生成模型’还是‘世界模拟器’)不能或不会成为智能体 AI。它们很可能成为,而且它们越好,这种可能性就越大,篡改数据并不是解决方案。

I do not have any solutions to this, other than to advise yet again to abandon the seductive, convenient, but wrong idea that tool AIs (under any branding, be it ‘tool AIs’ or ‘physics generative models’ or ‘world simulators’), cannot or will not be agent AIs. They may well be, and the better they get, the more likely it is, and tampering with data is not a solution.

1. [](https://gwern.net/scaling-hypothesis#fn1 "Link to footnote 1")

1. [](https://gwern.net/scaling-hypothesis#fn1 "Link to footnote 1")

鉴于论文算术基准上的评论数量,我应该指出,由于 BPE 编码问题,算术基准似乎大大低估了 GPT-3 的能力:例如,即使使用逗号也能显著提高其 5 位数加法能力。BPE 问题似乎也解释了在变位词/洗牌任务上的糟糕表现。对于任何需要字符级操作或理解的任务,这一点需要牢记。

Given the number of comments on the paper’s arithmetic benchmark, I should point out that the arithmetic benchmark appears to greatly understate GPT-3’s abilities due to the BPE encoding issue⁠: even using commas markedly improves its 5-digit addition ability, for example. The BPE issue also appears to explain much of the poor performance on the anagram/shuffling tasks. This is something to keep in mind for any task which requires character-level manipulation or understanding.[](https://gwern.net/scaling-hypothesis#fnref1)

2. [](https://gwern.net/scaling-hypothesis#fn2 "Link to footnote 2")

2. [](https://gwern.net/scaling-hypothesis#fn2 "Link to footnote 2")

关于隐式元学习,参见:Santoro 等人 2016 年/Wang 等人 2018 年(Botvinick 评论)/Botvinick 等人 2019a30061-0#deepmind)、Clune 2019 年、Schmidhuber 2015 年/2018 年、Weng 2018 年/Weng 2019 年。

On implicit ⁠meta-learning⁠, see: Santoro et al 2016⁠/Wang et al 2018⁠ (Botvinick commentary⁠)/Botvinick et al 2019a⁠30061-0#deepmind), Clune 2019⁠, Schmidhuber 2015⁠/2018⁠, Weng 2018⁠/Weng 2019⁠.[](https://gwern.net/scaling-hypothesis#fnref2)[](https://gwern.net/scaling-hypothesis#gwern-3956239092)

3. [](https://gwern.net/scaling-hypothesis#fn3 "Link to footnote 3")

3. [](https://gwern.net/scaling-hypothesis#fn3 "Link to footnote 3")

GPT-3 的算力成本不过几百万美元(截至 2020 年初),因为之前的广泛 Scaling 研究使得一次训练运行成为可能,并且运行成本低廉(第 39 页):“即使使用完整的 GPT-3 175B,从训练好的模型生成 100 页内容,成本大约为 0.4 千瓦时,或仅几美分的能源成本。”(同样,T5 也只训练了一次。)而且,对于一个模型的成本,GPT-3 API 用户已经证明,你可以获得相当于数百个较小的专用模型,每个模型都需要更多的研究人员、自定义数据集、无数次的训练运行和调整,假设这些模型能够被创建的话。(未来的口号:“一个模型,一个向量——一次。”)

GPT-3 hardly costs more than a few million dollars of compute (as of early 2020) as the extensive scaling research beforehand enabled one training run, and it is cheap to run (pg39): “Even with the full GPT-3 175B, generating 100 pages of content from a trained model can cost on the order of 0.4 kW-hr, or only a few cents in energy costs.” (Likewise, T5 was trained only once⁠.) And for the cost of one model, GPT-3 API users have shown that you get the equivalent of hundreds of smaller special-purpose models, each requiring more researchers, custom datasets, countless training runs, and tinkering, assuming said models could be created at all. (A slogan for the future: “One model, one vector—once.”)

相比之下,PDP-11 因其极低的成本而成为常见的学术主力机,仅需 110,099 美元(1970 年约 2 万美元),而第一台 Lisp Machine 的成本超过 258,956 美元(1972 年约 5 万美元)——对于工作站来说很贵,但与研究人员占用价值数千万美元的大型机相比,这很划算。IBM 的(否则无用的)深蓝 AI 项目据称在最终迭代中花费超过 12 美元(1997 年约 500 万美元)(有报道称 2.31 亿美元(1997 年约 1 亿美元)似乎与 Hsu 的《深蓝背后》第 187 页提到的宣传价值估计相混淆),而像 ITER 这样的大科学项目则花费超过 5000 倍的资金,大多以失败告终。(顺便说一句,粒子物理学家又回来要求超过 300 亿美元(2020 年约 240 亿美元),这大概是基于 LHC 超过 140 亿美元(2010 年约 90 亿美元)的投资所产生的科学革命和改变世界的突破,或者为(不)建造 SSC 而花费的 53 亿美元(1993 年约 20 亿美元)……)

For comparison, the PDP-11⁠ was a common academic workhorse due to its extremely low cost, a mere $110,099$20k 1970, while the first Lisp Machine⁠ cost >$258,956$50k 1972—expensive for a workstation but a bargain compared to researchers hogging mainframes costing tens of millions. IBM’s (otherwise useless) Deep Blue AI project reputedly cost >$12$5 1997 m for the final iteration (reports of $231$100 1997 m appear to be a confusion with the estimated value of _publicity_ mentioned in pg187 of Hsu’s _Behind Deep Blue_) and Big Science projects like ITER⁠ blow >5000× the funding to mostly fail. (The particle physicists, incidentally, are back asking for⁠ ≫$30$24 2020 b, based on, presumably the scientific revolutions & world-changing breakthroughs that the LHC’s >$14$9 2010 b investment produced, or the $5.30$2 1993 b spent to (not) build the SSC⁠…)

GPT-3 本可以在几十年前用全球计算资源和科学预算完成;用今天的硬件和预算,我们只是不知道或不愿意去做的事情,又能做什么呢?确实存在硬件过剩。(另见《全脑仿真路线图》和“2019 年 GPU 每 FLOPS 价格近期趋势”。)

GPT-3 could have been done decades ago with global computing resources & scientific budgets; what could be done with today’s hardware & budgets that we just don’t know or care to do? There _is_ a hardware overhang. (See also the _⁠Whole Brain Emulation Roadmap_⁠&“2019 recent trends in GPU price per FLOPS”⁠.)[](https://gwern.net/scaling-hypothesis#fnref3)

4. [](https://gwern.net/scaling-hypothesis#fn4 "Link to footnote 4")

4. [](https://gwern.net/scaling-hypothesis#fn4 "Link to footnote 4")

此外,由于训练与运行之间存在多个数量级的不对称性,神经网络本身也有额外的硬件过剩。迁移学习和元学习比基线模型训练快得多。你可以在没有任何梯度步骤的情况下‘训练’GPT-3——只需示例。你支付了极高的前期成本来训练一个统治一切的单一大型模型,然后以微小的边际成本在任何地方重复使用它。如果你训练了一个模型,那么一旦完成,你就会得到,除其他外:

Further, NNs have additional hardware overhangs of their own due to the many orders of magnitude asymmetry of training vs running. Transfer learning and meta-learning are so much faster than the baseline model training. You can ‘train’ GPT-3 without even any gradient steps—just examples. You pay the extremely steep upfront cost of One Big Model to Rule Them All, and then reuse it everywhere at tiny marginal cost. If you train a model, then as soon as it’s done you get, among other things:

* 在相同硬件上并行运行数千个副本的能力

* the ability to run thousands of copies in parallel on the same hardware

* 在像 AlphaGo 这样的环境中,我估计如果你重用相同的硬件仅运行带有原始模型精确副本的树搜索,可以获得数百个 ELO 强度的提升

* in a context like AlphaGo, I estimate several hundred ELO strength gains if you reuse the same hardware to merely run tree search with exact copies of the original model

* 对任何相关领域的元学习/迁移学习,将训练需求降低几个数量级

* meta-learning/transfer-learning to any related domain, cutting training requirements by orders of magnitude

* 模型压缩/蒸馏,以训练学生模型,其大小、FLOPS 或延迟仅为原始模型的一小部分(比例因任务、方法、领域、可接受的性能下降、目标硬件等而异,但通常极端,如 1/100)

* model compression/distillation to train student models which are a fraction of the size, FLOPS, or latency (ratios varying widely based on task, approach, domain, acceptable performance degradation, targeted hardware etc., but often extreme like 1⁄100 th)

* 在其他地方重用模型以立即增强其他模型(例如,为 DRL 智能体使用文本或图像嵌入)

* reuse of the model elsewhere to instantly power up other models (eg. use of text or image embeddings for a DRL agent)

* 边做边学/经验曲线效应(在信息技术中最高,在深度学习中也很高:Hernandez & Brown 2020),因此下一个从头开始的模型可能会便宜得多。

* learning-by-doing/experience curve effects⁠ (highest in information technologies, and high for DL: Hernandez & Brown 2020⁠), so the next from-scratch model may be much cheaper.

例如:在训练第一个 OpenAI Five(OA5)DoTA2 智能体时,经过所有迭代模型架构和游戏升级后,OA5 的第二次迭代“Rerun”是从头开始训练的。Rerun 只需要 20%的训练量就能达到“对 OpenAI Five 最终版本 98%的胜率”。正如作者所指出的:“理想的选择是从一开始就进行类似 Rerun 的训练,但这是不可能的——OpenAI Five 曲线代表了导致最终代码库、环境等的经验教训,没有这些,就不可能训练 Rerun。”

For example: after all the iterative model architecture & game upgrades done while training the first OpenAI Five⁠ (OA5) DoTA2 agent was completed, the second iteration of OA5, ⁠“Rerun”⁠, was trained from scratch. Rerun required only 20% of the training for a “98% win-rate against the final version of OpenAI Five.” As the authors note: “The ideal option would be to run Rerun-like training from the very start, but this is impossible—the OpenAI Five curve represents lessons learned that led to the final codebase, environment, etc., without which it would not be possible to train Rerun.”

* 通过消融和与原始模型比较,为工程更高效的模型提供基线

* baseline for engineering much more efficient ones by ablating and comparing with the original

5. [](https://gwern.net/scaling-hypothesis#fn5 "Link to footnote 5")

5. [](https://gwern.net/scaling-hypothesis#fn5 "Link to footnote 5")

例如,狭窄的上下文窗口严重限制了它,并激发了对高效注意力机制的需求。更广泛地说,GPT-3 没有做任何奇特的事情——没有使用大脑模仿学习或神经架构搜索来定制模型,没有在线超参数优化(可能超过 3 倍加速),甚至没有决定基本的超参数如宽度(正如 EfficientNet 所示,即使在“充分理解和手工优化的普通架构”中,这也能产生相当大的差异)。

eg. a narrow context window ⁠severely limits it⁠, and motivates the need for efficient attention⁠. More broadly, GPT-3 does nothing exotic—no use of brain imitation learning⁠ or neural architecture search to try to tailor the model, online hyperparameter optimization (possibly >3× speedup⁠) or even decide basic hyperparameters like widths (which as EfficientNet⁠ shows, can make quite a different even in “well-understood and hand-optimized vanilla architectures”).[](https://gwern.net/scaling-hypothesis#fnref5)

6. [](https://gwern.net/scaling-hypothesis#fn6 "Link to footnote 6")

6. [](https://gwern.net/scaling-hypothesis#fn6 "Link to footnote 6")

甚至没有 PDF——所以没有 Google Books、Arxiv、Libgen、Sci-Hub……

Not even PDFs—so no Google Books, no Arxiv, no Libgen, no Sci-Hub…[](https://gwern.net/scaling-hypothesis#fnref6)

7. [](https://gwern.net/scaling-hypothesis#fn7 "Link to footnote 7")

7. [](https://gwern.net/scaling-hypothesis#fn7 "Link to footnote 7")

从语言模型生成文本可以揭示知识的存在,但不能揭示其不存在,并且普遍认为当前粗糙的启发式方法如 top-k 不可能是最优的。

Generating text from a LM can reveal the presence of knowledge, but not its absence, and it is universally agreed that the current crude heuristic methods like top-_k_ cannot possibly be optimal.[](https://gwern.net/scaling-hypothesis#fnref7)

8. [](https://gwern.net/scaling-hypothesis#fn8 "Link to footnote 8")

8. [](https://gwern.net/scaling-hypothesis#fn8 "Link to footnote 8")

‘一个男人在医生办公室,医生告诉他:“我有好消息和坏消息要告诉你。”/男人说:“我现在听不了坏消息,先告诉我好消息吧。”/医生说:“好消息是你的阴茎有 18 英寸长。”/男人愣了一会儿,然后问:“坏消息是什么?”/医生说:“你的大脑在你的阴茎里。”’

‘A man is at the doctor’s office, and the doctor tells him, “I’ve got some good news and some bad news for you.” / The man says, “Well, I can’t take the bad news right now, so give me the good news first.” / The doctor says, “Well, the good news is that you have an 18-inch penis.” / The man looks stunned for a moment, and then asks, “What’s the bad news?” / The doctor says, “Your brain’s in your dick.”’[](https://gwern.net/scaling-hypothesis#fnref8)

9. [](https://gwern.net/scaling-hypothesis#fn9 "Link to footnote 9")

9. [](https://gwern.net/scaling-hypothesis#fn9 "Link to footnote 9")

特别是,样本效率随着模型大小的增加而提高,直到计算高效的 Scaling,而 GPT-2 可以在只看过一次数据后记住它——考虑到数据的长尾真实世界分布,这是一个理想的属性。(一个如何不写 Scaling 论文的例子是 Thompson 等人 2020 年,与前述论文形成鲜明对比——Thompson 等人根本没有提及这些论文!——他们试图不是从作者精心控制的实验中推断 Scaling,这些实验产生了非常紧密且高度可预测的曲线,而是从高度不同的研究论文中偶尔报告的数字来推断;不出所料,他们的曲线几乎无法预测任何东西,而且似乎无论如何都是严重高估。)

In particular, sample-efficiency increases with model size up to compute-efficient scaling, and GPT-2 can memorize data after seeing it only once⁠—a desirable property⁠ given long-tailed real-world distributions of data. (An example of how _not_ to do scaling papers is Thompson et al 2020⁠, which, in stark contrast to the foregoing papers—which Thompson et al do not mention at all!—attempts to infer scaling not from well-controlled experiments run by the authors, which yield extremely tight and highly predictive curves, but attempts to infer them from occasional reported numbers in highly disparate research papers; unsurprisingly, their curves barely predict anything and seem to be serious overestimates anyway.)

值得注意的是,追求大型模型几乎完全由 OpenAI 和行业实体驱动(后者满足于小得多的模型),而学术界表现出几乎完全的不感兴趣——甚至厌恶和愤怒,以及否认(有人可能会说“绿色 AI”是嫉妒得发绿)。尽管 Scaling 假设是‘显而易见的’,Scaling 是‘被预测的’,但真正去做的兴趣却少得惊人。也许我们应该更关注人们做了什么,而不是他们说了什么;并记住,成功的学术社区产生问题,而不是答案。

It is noteworthy that the pursuit of large models is driven almost exclusively by OpenAI & industry entities (the latter of which are content with far smaller models), and that academia has evinced an almost total disinterest—disgust & anger, even, and denial (one might say “green AI” is green with envy). For all that the scaling hypothesis is ‘obvious’ and scaling is ‘predicted’, there is remarkably little interest in actually _doing_ it. Perhaps we should pay more attention to what people do rather than what they say; and recall that successful academic communities produce questions, not answers.[](https://gwern.net/scaling-hypothesis#fnref9)

10. [](https://gwern.net/scaling-hypothesis#fn10 "Link to footnote 10")

10. [](https://gwern.net/scaling-hypothesis#fn10 "Link to footnote 10")

大致在 Chuan Li 的估计附近,使用名义标价而不考虑折扣(折扣可能很大,因为云计算的边际成本要低得多)。研发项目成本会高得多,但会分摊到所有后续模型和项目中。

Roughly around Chuan Li’s estimate, using nominal list prices without discounts (which could be steep as the marginal costs of cloud compute are substantially lower). The R&D project cost would be much higher, but is amortized over all subsequent models & projects.[](https://gwern.net/scaling-hypothesis#fnref10)

11. [](https://gwern.net/scaling-hypothesis#fn11 "Link to footnote 11")

11. [](https://gwern.net/scaling-hypothesis#fn11 "Link to footnote 11")

曼哈顿计划花费约 280 亿美元(1946 年约 20 亿美元)。

The Manhattan Project cost ~$28$2 1946 b.[](https://gwern.net/scaling-hypothesis#fnref11)

12. [](https://gwern.net/scaling-hypothesis#fn12 "Link to footnote 12")

12. [](https://gwern.net/scaling-hypothesis#fn12 "Link to footnote 12")

这让人想起顾客抱怨屠夫的笑话:

One is reminded of the joke about the customer complaining to the butcher:

“你的肉 10 美元一磅,而街对面的竞争对手只卖 1 美元!”“那就去买他的肉。”“我想买,但他没有。”“当我没有肉时,也只卖 1 美元。”

“Your meat is $10/lb, while your competitor across the street sells it at $1!” “So go buy his meat.” “I would, but he has none.” “When I don’t have any meat, it costs $1 too.”[](https://gwern.net/scaling-hypothesis#fnref12)

13. [](https://gwern.net/scaling-hypothesis#fn13 "Link to footnote 13")

13. [](https://gwern.net/scaling-hypothesis#fn13 "Link to footnote 13")

仿佛我们生活在一个只要足够努力,研究生就能用拉面预算登上月球的世界;仿佛在评估中只关注二氧化碳成本而不关注收益,就像做一把只有一片刀片的剪刀;或者仿佛试图不经过大型模型就创建小型模型的“绿色 AI”方法看起来越来越徒劳,像是在把好钱扔进坏钱里,并且是所有 AI 研究中最不绿色的……在某种程度上,所有前沿 AI 研究大约在 2010 年都可以用研究生级别的资金如 1,579 美元(2010 年约 1,000 美元)的硬件完成,而之前和之后几十年的 AI 研究都受益于大型计算机,这恰恰是对那个时代的控诉,表明那些研究是多么停滞不前的死胡同,其技术如此狭隘和受限,以至于无法从可用的大规模算力中受益。

As if we live in a world where grad students could go to the Moon on a ramen budget if we just wished hard enough, as if focusing on CO 2 costs & not benefits in our evaluations is not like making a scissor with only one blade, or as if “green AI” approaches to try to create small models without going through big models did not look increasingly futile and like throwing good money after bad, and were not the least green of all AI research… To the extent that all cutting-edge AI research ~2010 could be done with grad student money like $1,579$1k 2010 of hardware, where AI research in decades before & after benefited from big iron, that is an indictment of that era, demonstrating what a stagnant dead end that research was, that its techniques were so small-minded and hobbled it could not benefit from the available large-scale compute.[](https://gwern.net/scaling-hypothesis#fnref13)

14. [](https://gwern.net/scaling-hypothesis#fn14 "Link to footnote 14")

14. [](https://gwern.net/scaling-hypothesis#fn14 "Link to footnote 14")

有趣的事实:BiT 现在在预测(清理、校正后的)ImageNet 标签方面比原始 ImageNet 标签更准确。

Fun trivia: BiT is now more accurate⁠ at predicting (cleaned, corrected) ImageNet labels than the original ImageNet labels are.[](https://gwern.net/scaling-hypothesis#fnref14)

15. [](https://gwern.net/scaling-hypothesis#fn15 "Link to footnote 15")

15. [](https://gwern.net/scaling-hypothesis#fn15 "Link to footnote 15")

像 Dojolonga 等人 2020 年这样的图像 Scaling 实验的一个有趣方面是,即使原始任务的性能‘趋于平稳’并接近标签误差,迁移学习仍在继续改进。显然,内部表示,即使对于单纯的分类来说已经足够,因此分数只能增加很小的百分比,却变得更像人类——因为它编码了暗知识或更多的对抗鲁棒性?我注意到,对于语言模型,损失的最终分数似乎对生成样本质量有显著影响,也许是因为只有在所有更容易的建模完成后,懒惰的语言模型才被迫通过更正确地建模更复杂的事物(如逻辑、对象、世界知识等)来挤出下一部分性能。

One interesting aspect of image scaling experiments like Dojolonga et al 2020 is that even when performance is ‘plateauing’ on the original task & approaching label error, the transfer learning continues to improve. Apparently the internal representations, even when adequate for mere classification and so the score cannot increase more than a small percentage, become more human-like—because it’s encoding dark knowledge⁠ or more adversarial robustness⁠? I’ve noticed with language models, the final fractions of a loss appear to make a substantial difference to generated sample quality, perhaps because it is only after all the easier modeling is finished that the lazy language model is forced to squeeze out the next bit of performance by more correctly modeling more sophisticated things like logic, objects, world-knowledge, etc.[](https://gwern.net/scaling-hypothesis#fnref15)

16. [](https://gwern.net/scaling-hypothesis#fn16 "Link to footnote 16")

16. [](https://gwern.net/scaling-hypothesis#fn16 "Link to footnote 16")

这里的数字并不精确,仅用于说明;因为 BPE 不对应于任何直观概念,我将借用我观察字符 RNN 的经验,并讨论每个字符的损失而不是 BPE。

The numbers here are not exact and are for illustration; because BPEs don’t correspond to any intuitive, I am going to borrow from my observations watching char-RNNs, and talk about the loss per character instead of BPE.[](https://gwern.net/scaling-hypothesis#fnref16)

17. [](https://gwern.net/scaling-hypothesis#fn17 "Link to footnote 17")

17. [](https://gwern.net/scaling-hypothesis#fn17 "Link to footnote 17")

如果你看到数千张标记为‘狗’的图像和数千张标记为‘猫’的图像,你可以简单地学习单独的狗和猫分类器,而不必费心理解它们的共同方面,如被驯化的四足哺乳动物捕食者。如果你随后被要求分类‘雪貂’图像,这不会有帮助,但你没有被告知要这样做,所以这不是你的问题,因为如果你随后得到大量雪貂图像,你可以再学习另一个单独的雪貂分类器。

If you see thousands of images labeled ‘dog’ and thousands more labeled ‘cat’, you can simply learn separate dog & cat classifiers without bothering to understand their shared aspects like being domesticated quadruped mammal predators. This won’t be useful if you are then asked to classify ‘ferret’ images, but you weren’t asked to, so that’s not your problem, since you can just learn yet another separate classifier for ferrets if you then get a lot of ferret images.[](https://gwern.net/scaling-hypothesis#fnref17)

18. [](https://gwern.net/scaling-hypothesis#fn18 "Link to footnote 18")

18. [](https://gwern.net/scaling-hypothesis#fn18 "Link to footnote 18")

第 210-211 页,“安静的敌人”,《广岛的遗产》,Teller 1962 64ya。

pg210–211, “The Quiet Enemy”, _⁠The Legacy of Hiroshima_⁠, Teller 1962 64ya.[](https://gwern.net/scaling-hypothesis#fnref18)

19. [](https://gwern.net/scaling-hypothesis#fn19 "Link to footnote 19")

19. [](https://gwern.net/scaling-hypothesis#fn19 "Link to footnote 19")

解释关于 Transformer 实际上类似于 RNN 或实际上是 Hopfield 网络的各种论文的另一种方式是,将其视为表明,它们重要的不是与旧架构相比任何固有的新能力,而是某些更低层次的方面,比如在当代硬件上更高效的可训练性。

Another way of interpreting the various papers about how Transformers are actually like RNNs or are actually Hopfield networks⁠ is to take that as indicating that what is important about them is not any inherent new capability compared to older architectures, but some lower-level aspect like being more efficiently trainable on contemporary hardware.[](https://gwern.net/scaling-hypothesis#fnref19)

20. [](https://gwern.net/scaling-hypothesis#fn20 "Link to footnote 20")

20. [](https://gwern.net/scaling-hypothesis#fn20 "Link to footnote 20")

这些绝对预测性能与人类相比如何?很难说。人类/GPT-2/GPT-3 的困惑度可用基准似乎只有 WebText、Penn Tree Bank(PTB;基于 Brown 语料库)、1 Billion Word(1BW)和 LAMBADA。但覆盖范围参差不齐。

How do these absolute prediction performances compare to humans? It’s hard to say. The only available benchmarks for perplexity for humans/GPT-2/GPT-3 appear to be WebText, Penn Tree Bank⁠ (PTB; based on the Brown Corpus⁠), 1 Billion Word⁠ (1BW), and LAMBADA⁠. But coverage is spotty.

我没有找到 WebText 或 Penn Tree Bank 的人类基准,所以我无法比较人类与 GPT-2/GPT-3 的困惑度(GPT-2 PTB:35.7;GPT-3 PTB:20.5)。

I found no human benchmarks for WebText or Penn Tree Bank, so I can’t compare the human vs GPT-2/GPT-3 perplexities (⁠GPT-2 PTB⁠: 35.7; ⁠GPT-3 PTB⁠: 20.5).

GPT-2 在 1 Billion Word(1BW)基准上的困惑度为 43,而(高度外推的)人类困惑度为 12(有趣的是,使用 2012 年 14ya 的 LSTM RNN 外推,得出“还需要 10 到 20 年的研究才能达到人类性能”),但这可能是一个不公平的基准(“我们的模型仍然明显差于先前在 One Billion Word 基准上的工作(Chelba 等人 2013 年)。这可能是由于它既是最大的数据集,又具有一些最具破坏性的预处理——1BW 的句子级洗牌消除了所有长程结构。”),并且由于数据污染,1BW 已从 GPT-3 评估中删除(“我们省略了该工作中的 4 个与 Wikipedia 相关的任务,因为它们完全包含在我们的训练数据中,并且由于数据集中很大一部分包含在我们的训练集中,我们也省略了十亿词基准。”)。

⁠GPT-2⁠ was benchmarked at 43 perplexity on the 1 Billion Word (1BW) benchmark vs a (highly extrapolated) human perplexity of 12⁠ (which interestingly extrapolates, using 2012 14ya LSTM RNNs, that “10 to 20 more years of research before human performance is reached”), but that may be an unfair benchmark (“Our model is still substantially worse than prior work on the One Billion Word Benchmark (⁠Chelba et al 2013⁠). This is likely due to a combination of it being both the largest dataset and having some of the most destructive pre-processing—1BW’s sentence level shuffling removes all long-range structure.”) and 1BW was dropped from the GPT-3 evaluation due to data contamination (“We omit the 4 Wikipedia-related tasks in that work because they are entirely contained in our training data, and we also omit the one-billion word benchmark due to a high fraction of the dataset being contained in our training set.”).

LAMBADA 的基准困惑度:GPT-2 为 8.6,GPT-3 为零样本 3.0/少样本 1.92。OA 在其 GPT-2 博客文章(而非论文)中声称人类困惑度为 1-2,但没有提供来源,我也找不到任何来源。(作者可能是根据 LAMBADA 的构建方式猜测的:示例通过两个独立的人类评分者是否提供相同的正确答案来过滤,这为人类预测答案的能力设定了下限。)

LAMBADA was benchmarked at a ⁠GPT-2 perplexity⁠ of 8.6, and a ⁠GPT-3 perplexity⁠ of 3.0 (zero-shot) / 1.92 (few-shot). ⁠OA claims⁠ in their GPT-2 blog post (but not the paper) that human perplexity is 1–2, but provides no sources and I couldn’t find any. (The authors might be guessing based on how LAMBADA was constructed: examples were filtered by whether two independent human raters provided the same right answer, which lower bounds how good humans must be at predicting the answer.)

因此,总体而言,最佳猜测似乎是 GPT-3 的绝对误差大约是人类的两倍。这意味着需要大量(但远非不可能)的算力才能根据当前的缩放定律完全弥合剩余的差距。如果我们不负责任地进一步外推 WebText 的 Scaling 曲线,假设 GPT-3 在其当前 WebText 困惑度 1.73 下具有两倍于人类的误差(因此人类约为 0.86),那么我们需要 2.57 ⋅ (3.64 ⋅ (10^3 ⋅ _x_))^-0.048 = 0.86,其中_x_ = 2.2e6,即 GPT-3 算力的 2,200,000 倍。(这大致相当于美国入侵伊拉克的成本。)

So overall, it looks like the best guess is that GPT-3 continues to have somewhere around twice the absolute error of a human. This implies it will take a large (yet, far from impossible) amount of compute to fully close the remaining gap with the current scaling laws. If we irresponsibly extrapolate out the WebText scaling curve further, assume GPT-3 has twice the error of a human at its current WebText perplexity of 1.73 (and so humans are ~0.86), then we need 2.57 ⋅ (3.64 ⋅ (10 3 ⋅ _x_))-0.048 = 0.86, where _x_ = 2.2e6 or 2,200,000× the compute of GPT-3. (This would roughly equal the cost to the USA of invading Iraq.)

如果我们想象峰值 AI 算力使用每 3.4 个月翻一番,那么 2.2e6 将是 22 次翻倍——即 6.3 年,在 2027 年。大多数人认为这种算力趋势很快就会崩溃,这种预测就是一个很好的理由!

If we imagine that ⁠peak AI compute usage doubles every 3.4 months⁠, then 2.2e6 would be 22 doublings away—or 6.3 years, in 2027. Most people believe that that compute trend must break down soon, and this sort of prediction is a good reason why!

另一方面,Hernandez & Brown 2020 的估计是,扣除硬件和算法进步,固定性能水平的成本每 16 个月减半;因此,如果 GPT-3 在 2020 年初花费约 629 万美元(2020 年约 500 万美元),那么到 2021 年中,它将花费 314 万美元(2020 年约 250 万美元),依此类推。类似地,一个需要 2.2e6 倍算力的 GPT 人类在 2020 年将花费约 13 万亿美元(2020 年约 10 万亿美元),但在 14 次减半(18 年)后,到 2038 年将花费约 12.6 亿美元(2020 年约 10 亿美元)。

Going the other direction, ⁠Hernandez & Brown 2020’s⁠ estimate is that, net of hardware & algorithmic progress, the cost of a fixed level of performance halves every 16 months; so if GPT-3 cost ~$6.29$5 2020 m in early 2020, then it’ll cost $3.14$2.50 2020 m around mid-2021, and so on. Similarly, a GPT-human requiring 2.2e6× more compute would presumably cost on the order of $13$10 2020 trillion in 2020, but after 14 halvings (18 years) would cost $1.26$1 2020 b in 2038.[](https://gwern.net/scaling-hypothesis#fnref20)[](https://gwern.net/scaling-hypothesis#gwern-1230483350)

21. [](https://gwern.net/scaling-hypothesis#fn21 "Link to footnote 21")

21. [](https://gwern.net/scaling-hypothesis#fn21 "Link to footnote 21")

截至 2020 年 12 月,半年后,几乎没有研究人员愿意公开记录他们预测未来 1t、10t 或 100t 模型将具有或不具有哪些具体能力,以及在什么规模下哪些缺失的能力将出现——就像没有人成功预测 GPT-2 或 GPT-3 的具体能力一样。

As of December 2020, half a year later, almost no researcher has been willing to go on record as saying what specific capabilities they predict future 1t, 10t, or 100t models will have or not have, and at what size which missing capabilities will emerge—just as no one is on record successfully predicting GPT-2 or GPT-3’s specific capabilities.[](https://gwern.net/scaling-hypothesis#fnref21)

22. [](https://gwern.net/scaling-hypothesis#fn22 "Link to footnote 22")

22. [](https://gwern.net/scaling-hypothesis#fn22 "Link to footnote 22")

另见 Sutskever 的 DRL 演讲,以及 Wojciech Zaremba 关于 OA5 的评论(转录):

See also ⁠Sutskever’s DRL talk⁠, and Wojciech Zaremba’s⁠⁠comments about OA5⁠ (⁠transcript⁠):

23. [](https://gwern.net/scaling-hypothesis#fn23 "Link to footnote 23")

23. [](https://gwern.net/scaling-hypothesis#fn23 "Link to footnote 23")

生产服务,尤其是免费生产服务,通常远远落后于最前沿实验室中未发表的 SOTA。当然,后者才是预测 AI 进展或 AI 风险的唯一重要因素,但人们会坚持用奇怪的指标来衡量 AI 进展,比如任意免费服务去年能做什么。作为经验法则,假设:如果你使用的是无需登录的免费服务,质量至少落后 SOTA 2 年;需要登录的免费服务,超过 1.5 年;付费服务,超过 1 年;最近发布的研究论文,超过 6 个月。

Production services, especially _free_ production services, usually lag long after the unpublished SOTA inside the most cutting-edge lab. The second is the only thing that matters for predicting AI progress or AI risk, of course, but people will insist on measuring AI progress by bizarre metrics like what an arbitrary free service could do last year. As a rule of thumb, assume that: if you are using a free service with no login, the quality is _at least_ 2 years behind SOTA; free with a login, >1.5 years; paid service, >1 year; & recently-released research paper, >6 months.[](https://gwern.net/scaling-hypothesis#fnref23)

24. [](https://gwern.net/scaling-hypothesis#fn24 "Link to footnote 24")

24. [](https://gwern.net/scaling-hypothesis#fn24 "Link to footnote 24")

特别是 Demis Hassabis;我不确定 Shane Legg 目前的观点,尽管考虑到他 2009 年创立 DeepMind 时预测的准确性以及他 2018 年的评论,他可能没有太大改变他的观点,即 AI 将由(已实现的)指数级算力增长赋能,或者他关于 AGI 的预测大约在 2028 年。(这与最新的 Metaculus 预测一致。)

Particularly Demis Hassabis⁠; I’m not sure about Shane Legg’s⁠current views⁠, although given the accuracy of his 2009 predictions⁠ while founding DeepMind & his 2018 comments⁠, he probably hasn’t much changed his views that AI will be empowered by the (realized) exponential compute gains or his AGI forecast of ~2028⁠. (This is consistent with the latest Metaculus⁠forecasts⁠.)[](https://gwern.net/scaling-hypothesis#fnref24)

25. [](https://gwern.net/scaling-hypothesis#fn25 "Link to footnote 25")

25. [](https://gwern.net/scaling-hypothesis#fn25 "Link to footnote 25")

当面临选择:要么承认他们所有花哨的艰苦工作都是死胡同,吞下苦涩的教训,开始预算数千万的算力;要么写一条轻蔑的推文解释“实际上,GPT-3 表明 Scaling 是死胡同,这是一场环境灾难,而且它只是模仿智能”——大多数人会忙着写推文。

When faced with the choice between having to admit all their fancy hard work is a dead-end, swallow the bitter lesson, and start budgeting tens of millions of compute, or instead writing a disdainful tweet explaining how, “_actually_, GPT-3 shows that scaling is a dead end, it’s an environmental catastrophe, and it’s just imitation intelligence anyway”—most people will get busy on the tweet

26. [](https://gwern.net/scaling-hypothesis#fn26 "Link to footnote 26")

26. [](https://gwern.net/scaling-hypothesis#fn26 "Link to footnote 26")

像 GShard 这样的混合专家模型或像 DynamicEmbedding 这样的嵌入与 GPT-3 这样的‘密集’模型不可比,因为在某种意义上,训练具有数十亿‘参数’的模型一直都很便宜且容易,比如极大的嵌入;然而,这些参数作用不大,更像是一堆浅层模型背靠背粘在一起。它们可能不会学到具有相同名义参数计数的密集模型会学到的有趣东西。

A mixture-of-expert model like GShard or an embedding like DynamicEmbedding is not comparable to ‘dense’ models like GPT-3, as it’s always been cheap & easy to train models with billions of ‘parameters’ in some sense, like extremely large embeddings; however, these parameters do little, and are more like a few hundred shallow models glued back-to-back. They probably do not learn the same interesting things that a dense model would with the same nominal parameter count.[](https://gwern.net/scaling-hypothesis#fnref26)

27. [](https://gwern.net/scaling-hypothesis#fn27 "Link to footnote 27")

27. [](https://gwern.net/scaling-hypothesis#fn27 "Link to footnote 27")

这似乎是评论者的一个盲点:假设如果必要的资源存在,那么它就会被使用。例如,Jim Gray(2007 年去世,19ya)在 1999 年 6 月(27ya)对图灵的连接主义硬件论证略加嘲讽,指出(使用人类大脑计算能力的乐观下限):

This seems to be a bit of a blind spot by commentators: the assumption that if the necessary resource _exists_, then it will be _used_. For example, Jim Gray (d. 2007 19ya) in June 1999 27ya⁠pokes a bit of fun⁠ at Turing’s connectionist hardware argument by noting that (using an optimistic lower bound on human brain computational power):

事后看来,我们可以说,1999 年(27ya)的超级计算机确实可以展示出比它们当时所展示的更令人印象深刻的智能水平,而且 1999 年(27ya)在超级计算机上运行的软件也永远不会导致有意义的 AI 进展,这之间没有特别的矛盾或神秘——只是没有人尝试。没有超级计算机所有者会允许它被占用数年,进行使连接主义方法如 RNN 或 CNN 工作所需的小而关键的迭代。因此,确实需要一些完全不同且激进的东西——但我们已经知道解决方案是什么样子。

With the benefit of hindsight, we can say that it is true that supercomputers in 1999 27ya could have been showing far more impressive levels of intelligence than they were, and that it was also true that the software being run on the supercomputers in 1999 27ya were never going to lead to meaningful AI progress, and that there is no particular contradiction or mystery—it was simply that no one was trying. No supercomputer owner was going to let it be tied up for years doing the minor-yet-critical iteration to make connectionist approaches like RNNs or CNNs work. Thus, something quite different & radical was indeed needed—but we already knew what the solution looked like.[](https://gwern.net/scaling-hypothesis#fnref27)

28. [](https://gwern.net/scaling-hypothesis#fn28 "Link to footnote 28")

28. [](https://gwern.net/scaling-hypothesis#fn28 "Link to footnote 28")

引人注目的是,截至 2020 年,这仍然成立:例如,我在 Summit 上看到的唯一深度学习研究是材料科学和生物学。(在再次检查 Arxiv 时,我确实找到了一篇使用 Summit 资源的非 STEM 论文:Lin 等人 2019 年,专注于训练视频分类模型中的系统工程。)

Strikingly, as of 2020, this is _still_ true: eg. the only deep learning research I have seen done on Summit⁠) were materials⁠science⁠&biology⁠. (In double-checking Arxiv, I did find one non-STEM paper using Summit resources: Lin et al 2019⁠, focusing on systems engineering in training a video classification model.)[](https://gwern.net/scaling-hypothesis#fnref28)

29. [](https://gwern.net/scaling-hypothesis#fn29 "Link to footnote 29")

29. [](https://gwern.net/scaling-hypothesis#fn29 "Link to footnote 29")

Peter Norvig 提供了一个例子,说明当研究生负担不起使神经网络工作所需的计算能力时会发生什么:

Peter Norvig⁠⁠offers an example⁠ of what happens when grad students _can’t_ afford the necessary computing power to make neural nets work:

据我估计,Norvig 的尝试使用了相当于当代 GPU 时间 0.8 毫秒的计算量。

By my estimate, Norvig’s attempt used the equivalent of 0.8 _milliseconds_ of contemporary GPU-time.

(大约在 1981 年,一台昂贵的 PC,成本相当于超过 6,289 美元(2020 年约 5,000 美元),这种 PC 可能分配给研究生每人一台,可能配备额外的 Intel 8087 浮点协处理器,能够进行 50,000 FP64 FLOPS;保守假设‘过夜’+‘再多一天’≤2 天,那么 Norvig 的实验使用了 2 天×24 小时×60 分钟×60 秒×50,000 = 8×10^9 FLOPS;2020 年的 Nvidia A100 GPU 名义价格约为 12,578 美元(2020 年约 10,000 美元),拥有 9.7 FP64 TFLOPS 或 9,700,000,000,000 FLOPS(在更有用的低精度模式如 FP32 中更多,但 1981 年 45ya 的 ML 不知道这一点);因此,8×10^9 / 9.7×10^12 = 8×10^-4 秒 = 0.8 毫秒。)

(In ~1981, an expensive PC costing the equivalent of >$6,289$5k 2020, of the sort a high-powered AI lab might allocate 1 apiece to grad students, might have an additional Intel 8087⁠ floating-point coprocessor⁠ capable of 50,000 FP64 FLOPS; conservatively assuming that ‘overnight’ + ‘one more day’ ≤ 2 days, then Norvig’s experiment used 2d × 24h × 60m × 60s × 50,000 = 8×10 9 FLOPS; a 2020 Nvidia A100⁠) GPU nominally priced ~$12,578$10k 2020 boasts 9.7 FP64 TFLOPS or 9,700,000,000,000 FLOPS (and far more in the more useful low-precision regimes like FP32, but 1981 45ya ML didn’t know that); thus, 8 9 / 9.7×10 12 = 8×10−4 seconds = 0.8 milliseconds.)[](https://gwern.net/scaling-hypothesis#fnref29)

30. [](https://gwern.net/scaling-hypothesis#fn30 "Link to footnote 30")

30. [](https://gwern.net/scaling-hypothesis#fn30 "Link to footnote 30")

Jeff Dean 指出,“也许不幸的是,就在我们开始拥有足够的计算性能来应对有趣的现实世界问题,并且机器学习规模的扩大和应用范围的增加导致对额外计算资源解决更大问题的巨大需求时,整个计算行业在通用 CPU 性能的年同比改进方面经历了急剧放缓。”在计算观点下,这并非巧合:算力,而非算法,是关键因素;生物系统通常在一个数量级或更少范围内接近任务的理论最优值;并且越接近最优,进展越慢;因此,当人工计算终于开始处理“有趣的现实世界问题”时,它必然接近其极限。(情况本可以不同:摩尔定律本可以在生物效率的许多数量级之前停止,或者超过许多数量级,没有时间上的巧合,而 AI 因其他原因发生。)

Jeff Dean⁠ notes, “It is perhaps unfortunate that just as we started to have enough computational performance to start to tackle interesting real-world problems and the increased scale and applicability of machine learning has led to a dramatic thirst for additional computational resources to tackle larger problems, the computing industry as a whole has experienced a dramatic slowdown in the year-over-year improvement of general purpose CPU performance.” Under the computational view, this is not a coincidence: compute, not algorithms, are the critical factor; biological systems often come within orders of magnitude, or less, of the theoretical optimum for a task; and the closer one comes to optimal, the slower progress becomes; so, just as artificial computation finally starts doing “interesting real-world problems”, it necessarily is approaching its limits. (It could have been otherwise: Moore’s law could have stopped short by many orders of magnitude of biological efficiency, or surpassed it by many orders, with no temporal coincidence, and AI happened for other reasons.)[](https://gwern.net/scaling-hypothesis#fnref30)

31. [](https://gwern.net/scaling-hypothesis#fn31 "Link to footnote 31")

31. [](https://gwern.net/scaling-hypothesis#fn31 "Link to footnote 31")

既然 GPT-3 的少样本学习和 T5 微调已经开始让像 Gary Marcus 这样的人对 WinoGrande 感到些许不安,他们已经开始准备借口,解释为什么 Winograd 模式并不是真正衡量常识推理/智能的好方法(因为智能,当然,是 AI 还不能做的任何事情)。

Now that GPT-3’s few-shot and T5 finetuning⁠ have begun to make people like Gary Marcus feel slightly nervous about WinoGrande, they have begun preparing⁠their excuses⁠ for why Winograd schemas weren’t _really_⁠ good measures of commonsense reasoning/intelligence (because intelligence, of course, is whatever AI can’t do yet).[](https://gwern.net/scaling-hypothesis#fnref31)

32. [](https://gwern.net/scaling-hypothesis#fn32 "Link to footnote 32")

32. [](https://gwern.net/scaling-hypothesis#fn32 "Link to footnote 32")

Feynman:“有几次提到以前的飞行;这些飞行的接受和成功被视为安全的证据。但侵蚀和泄漏并不是设计所预期的。它们是事情出错的警告。设备没有按预期运行,因此存在危险,它可能以意想不到且未被完全理解的方式出现更大的偏差。这一危险以前没有导致灾难,并不能保证下次不会,除非它被完全理解。”

Feynman⁠: “There are several references to previous flights; the acceptance and success of these flights are taken as evidence of safety. But erosion and blowby are not what the design expected. They are warnings that something is wrong. The equipment is not operating as expected, and therefore there is a danger that it can operate with even wider deviations in the unexpected and not thoroughly understood way. The fact that this danger did not lead to catastrophe before is no guarantee that it will not the next time, unless it is completely understood.”

这让人想起中国皇帝和 Shaka Zulu 对西方技术的灾难性轻视:错误不在于轻视技术在实际上的重要性(可以说它确实不重要),也不在于未能意识到他们在几千年最重要的地缘政治发展中已经落后了几个世纪,而在于未能承认他们无法解释为什么西方技术变得如此之快如此之好,因此无法知道它不会变得更好。

One is reminded of the catastrophic dismissals of Western technology by the Chinese emperors &Shaka Zulu⁠: the error was not in dismissing the technology as practically unimportant (it arguably was), nor in failing to realize that they were already centuries behind in the most important geopolitical development in millennia, but in failing to acknowledge that they couldn’t explain why the Western technology had gotten so good so fast and thus couldn’t know that it wouldn’t get much better still.[](https://gwern.net/scaling-hypothesis#fnref32)

33. [](https://gwern.net/scaling-hypothesis#fn33 "Link to footnote 33")

33. [](https://gwern.net/scaling-hypothesis#fn33 "Link to footnote 33")

正如我最喜欢的维特根斯坦轶事所说,日心说让每个人都觉得是假的,因为事物看起来并不像地球以天文速度围绕一颗恒星旋转,而是地球完全静止,其他一切围绕它旋转(Anscombe 1963 63ya,《维特根斯坦〈逻辑哲学论〉导论》):

As my favorite Wittgenstein anecdote goes, heliocentrism strikes everyone as false because things just don’t _look_ as if the Earth whirls at astronomical velocities around a star, but as if the Earth is perfectly still and everything else whirls around it (Anscombe 1963 63ya, _An Introduction to Wittgenstein’s Tractatus_):

34. [](https://gwern.net/scaling-hypothesis#fn34 "Link to footnote 34")

34. [](https://gwern.net/scaling-hypothesis#fn34 "Link to footnote 34")

这种联系不仅仅是表面的——很多强化学习工作借鉴了物理学和变分原理的形式类比。

This connection is more than superficial—a lot of RL work draws on formal analogies to physics and variational principles.[](https://gwern.net/scaling-hypothesis#fnref34)

反向链接 Backlinks(https://gwern.net/scaling-hypothesis#backlinks-section "Link to section: § 'Backlinks'")

* OpenAI 联合创始人 Sutskever 新成立的安全导向 AI 初创公司 SSI 筹集 10 亿美元⁠:

* OpenAI co-founder Sutskever’s new safety-focused AI startup SSI raises $1 billion⁠:

* 基础模型能否推理因果关系?⁠:

* Can Foundation Models Talk Causality?⁠:

* GPT-3 创意小说⁠(⁠完整上下文⁠):

* GPT-3 Creative Fiction⁠ (⁠full context⁠):

* ‘小团体’目录⁠(⁠完整上下文⁠):

* ‘small groups’ directory⁠ (⁠full context⁠):

* ARPA 和 SCI:冲浪 AI⁠(⁠完整上下文⁠):

* ARPA and SCI: Surfing AI⁠ (⁠full context⁠):

* 长期记忆:信息规模与大脑尺寸的缩放⁠:

* Long-term memory: scaling of information to brain size⁠:

* 迈向基准测试 LLM 多样性与创造力⁠(⁠完整上下文⁠):

* Towards Benchmarking LLM Diversity & Creativity⁠ (⁠full context⁠):

* 绝对单元神经网络:基于回归的 MLP 处理一切⁠(⁠完整上下文⁠):

* Absolute Unit NNs: Regression-Based MLPs for Everything⁠ (⁠full context⁠):

* 缩放 MLP:归纳偏置的故事⁠:

* Scaling MLPs: A Tale of Inductive Bias⁠:

* 用于上传的模块化大脑 AUNN⁠(⁠完整上下文⁠):

* Modular Brain AUNNs for Uploads⁠ (⁠full context⁠):

* RL 智能体的自由发挥阶段⁠(⁠完整上下文⁠):

* Free-Play Periods for RL Agents⁠ (⁠full context⁠):

* GAN 没有失败,它们被抛弃了⁠(⁠完整上下文⁠):

* GANs Didn’t Fail, They Were Abandoned⁠ (⁠full context⁠):

* 思维链提示引发大型语言模型的推理⁠:

* Chain-of-Thought Prompting Elicits Reasoning in Large Language Models⁠:

* Grokking:在小算法数据集上超越过拟合的泛化⁠:

* Grokking: Generalization Beyond Overfitting On Small Algorithmic Datasets⁠:

* 用于音乐生成的 GPT-2 偏好学习⁠(⁠完整上下文⁠):

* GPT-2 Preference Learning for Music Generation⁠ (⁠full context⁠):

* ‘神经网络稀疏性’目录⁠(⁠完整上下文⁠):

* ‘NN sparsity’ directory⁠ (⁠full context⁠):

* ‘AI 缩放’目录⁠(⁠完整上下文⁠):

* ‘AI scaling’ directory⁠ (⁠full context⁠):

* ARPA 和 SCI:冲浪 AI⁠(⁠完整上下文⁠):

* ARPA and SCI: Surfing AI⁠ (⁠full context⁠):

* GPT-3 创意小说⁠(⁠完整上下文⁠):

* GPT-3 Creative Fiction⁠ (⁠full context⁠):

* GPT-3 创意小说⁠(⁠完整上下文⁠):

* GPT-3 Creative Fiction⁠ (⁠full context⁠):

类似链接 Similar Links(https://gwern.net/scaling-hypothesis#similars-section "Link to section: § 'Similar Links'")

* DALL·E 1:从文本生成图像:我们训练了一个名为 DALL·E 的神经网络,它能够根据自然语言表达的广泛概念从文本描述生成图像

* DALL·E 1: Creating Images from Text: We’ve trained a neural network called DALL·E that creates images from text captions for a wide range of concepts expressible in natural language⁠

* 大型生成模型中的可预测性与惊喜

* Predictability and Surprise in Large Generative Models⁠

* 语言模型 Scaling:训练 Gopher 的方法、分析与见解

* Scaling Language Models: Methods, Analysis & Insights from Training Gopher⁠

* CT0:微调语言模型是持续学习者

* CT0: Fine-tuned Language Models are Continual Learners⁠

* 神经缩放定律的一个可解模型

* A Solvable Model of Neural Scaling Laws⁠

* 神经语言模型的缩放定律

* Scaling Laws for Neural Language Models⁠

* 关于基础模型的机遇与风险

* On the Opportunities and Risks of Foundation Models⁠

参考文献 Bibliography(https://gwern.net/scaling-hypothesis#link-bibliography-section "Link to section: § 'Bibliography'")

[⁠[页面中使用的链接/参考文献目录]⁠](https://gwern.net/metadata/annotation/link-bibliography/%252Fscaling-hypothesis.html)

[⁠[Bibliography of links/references used in page]⁠](https://gwern.net/metadata/annotation/link-bibliography/%252Fscaling-hypothesis.html)

[[发送匿名反馈]](https://docs.google.com/forms/d/e/1FAIpQLSd7uqL7B_l1HFIfXc8D_nZyumaOv58msK7jhl4XzQjWODWKdA/viewform "用于向 Gwern Branwen 提交匿名反馈的 Google Docs 网页表单")

[[Send Anonymous Feedback]](https://docs.google.com/forms/d/e/1FAIpQLSd7uqL7B_l1HFIfXc8D_nZyumaOv58msK7jhl4XzQjWODWKdA/viewform "Google Docs web form for submitting anonymized feedback to Gwern Branwen")

[[每日名言]](https://gwern.net/metadata/today-quote.html "内容正在加载,请稍候。")

[[Quote Of The Day]](https://gwern.net/metadata/today-quote.html "Content is loading. Please wait.")

[[每日站点]](https://gwern.net/metadata/today-site.html "内容正在加载,请稍候。")

[[Site Of The Day]](https://gwern.net/metadata/today-site.html "Content is loading. Please wait.")

[[每日注释]](https://gwern.net/metadata/today-annotation.html "内容正在加载,请稍候。")

[[Annotation Of The Day]](https://gwern.net/metadata/today-annotation.html "Content is loading. Please wait.")

[[广告拦截公益公告]](https://gwern.net/metadata/psa-adblock.html "内容正在加载,请稍候。")

[[adblock public service announcement]](https://gwern.net/metadata/psa-adblock.html "Content is loading. Please wait.")

互动版:图/公式 + 针对本篇提问 →