Situational Awareness I: From GPT-4 to AGI
打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→第一部分:从 GPT-4 到 AGI:计算数量级提升。第二部分:从 AGI 到超级智能:智能爆炸。第三部分 a:竞逐万亿美元集群。
* I. From GPT-4 to AGI: Counting the OOMs * II. From AGI to Superintelligence: the Intelligence Explosion * IIIa. Racing to the Trillion-Dollar Cluster
* I. 从 GPT-4 到 AGI:计算 OOM 数量
* I. From GPT-4 to AGI: Counting the OOMs
* II. 从 AGI 到超级智能:智能爆炸
* II. From AGI to Superintelligence: the Intelligence Explosion
* IIIa. 竞相建造万亿美元集群
* IIIa. Racing to the Trillion-Dollar Cluster
* IIIb. 封锁实验室:AGI 的安全保障
* IIIb. Lock Down the Labs: Security for AGI
到 2027 年实现 AGI 的可能性惊人地高。GPT-2 到 GPT-4 在 4 年内将能力从大约学龄前儿童水平提升到了聪明高中生水平。追踪算力(每年约 0.5 个数量级或 OOM)、算法效率(每年约 0.5 个 OOM)以及“去束缚”增益(从聊天机器人到智能体)的趋势线,我们预计到 2027 年将再次出现类似从学龄前儿童到高中生的质的飞跃。
AGI by 2027 is strikingly plausible. GPT-2 to GPT-4 took us from ~preschooler to ~smart high-schooler abilities in 4 years. Tracing trendlines in compute (~0.5 orders of magnitude or OOMs/year), algorithmic efficiencies (~0.5 OOMs/year), and “unhobbling” gains (from chatbot to agent), we should expect another preschooler-to-high-schooler-sized qualitative jump by 2027.
* 附录。快速穿越 OOM:要么这十年,要么永远没戏
* Addendum. Racing through the OOMs: It’s this decade or bust
GPT-4 的能力让许多人震惊:一个能编写代码和文章、能推理困难数学问题并通过大学考试的人工智能系统。几年前,大多数人认为这些是不可逾越的障碍。
GPT-4’s capabilities came as a shock to many: an AI system that could write code and essays, could reason through difficult math problems, and ace college exams. A few years ago, most thought these were impenetrable walls.
但 GPT-4 只是深度学习十年飞速发展的延续。十年前,模型几乎无法识别简单的猫狗图片;四年前,GPT-2 几乎无法拼凑出稍微合理的句子。现在我们正迅速饱和所有我们能想到的基准测试。然而,这一戏剧性的进步仅仅是深度学习规模扩张的持续趋势的结果。
But GPT-4 was merely the continuation of a decade of breakneck progress in deep learning. A decade earlier, models could barely identify simple images of cats and dogs; four years earlier, GPT-2 could barely string together semi-plausible sentences. Now we are rapidly saturating all the benchmarks we can come up with. And yet this dramatic progress has merely been the result of consistent trends in scaling up deep learning.
有些人更早看到了这一点。他们曾被嘲笑,但他们所做的只是相信趋势线。趋势线是强烈的,他们是对的。模型只是想学习;你扩大它们,它们就学得更多。
There have been people who have seen this for far longer. They were scoffed at, but all they did was trust the trendlines. The trendlines are intense, and they were right. The models, they just want to learn; you scale them up, and they learn more.
我提出以下主张:到 2027 年,模型能够完成 AI 研究员/工程师的工作,这是惊人地合理的。这不需要相信科幻小说;只需要相信图表上的直线。
I make the following claim: it is strikingly plausible that by 2027, models will be able to do the work of an AI researcher/engineer. That doesn’t require believing in sci-fi; it just requires believing in straight lines on a graph.
_基于本文讨论的公开估计,过去和未来有效算力(包括物理算力和算法效率)规模扩张的粗略估计。随着我们扩大模型,它们持续变得更聪明,通过“数 OOM”,我们可以粗略了解(近期)未来应预期的模型智能水平。(此图仅显示基础模型的规模扩张;“去束缚”未显示。)_
_Rough estimates of past and future scaleup of effective compute (both physical compute and algorithmic efficiencies), based on the public estimates discussed in this piece. As we scale models, they consistently get smarter, and by “counting the OOMs” we get a rough sense of what model intelligence we should expect in the (near) future. (This graph shows only the scaleup in base models; “unhobblings” are not pictured.)_
在本文中,我将简单地“数 OOM”(OOM = 数量级,10 倍 = 1 个数量级):观察 1)算力、2)算法效率(可视为增长“有效算力”的算法进步)和 3)“去束缚”增益(修复模型默认情况下被束缚的明显方式,解锁潜在能力并赋予工具,导致有用性的阶跃变化)的趋势。我们追踪 GPT-4 之前四年内每项的增长,以及之后四年(到 2027 年底)应预期的增长。鉴于深度学习在有效算力每增加一个 OOM 时的一致改进,我们可以用此来预测未来的进展。
In this piece, I will simply “count the OOMs” (OOM = order of magnitude, 10x = 1 order of magnitude): look at the trends in 1) _compute_, 2) _algorithmic efficiencies_ (algorithmic progress that we can think of as growing “effective compute”), and 3) _”unhobbling” gains_ (fixing obvious ways in which models are hobbled by default, unlocking latent capabilities and giving them tools, leading to step-changes in usefulness). We trace the growth in each over four years before GPT-4, and what we should expect in the four years after, through the end of 2027. Given deep learning’s consistent improvements for every OOM of effective compute, we can use this to project future progress.
公开地,自 GPT-4 发布以来的一年里,事情一直很平静,因为下一代模型正在酝酿中——这导致一些人宣称停滞不前,深度学习正在碰壁。
Publicly, things have been quiet for a year since the GPT-4 release, as the next generation of models has been in the oven—leading some to proclaim stagnation and that deep learning is hitting a wall.
1 但通过数 OOM,我们可以一窥实际应预期的情况。
1 But by counting the OOMs, we get a peek at what we should actually expect.
结论相当简单。从 GPT-2 到 GPT-4——从有时能拼凑出几个连贯句子的模型,到通过高中考试的模型——并非一次性的收益。我们正极其迅速地穿越 OOM,数字表明我们应预期在四年内再次实现约 100,000 倍的有效算力规模扩张——导致另一次 GPT-2 到 GPT-4 规模的质的飞跃。此外,关键的是,这不仅仅意味着更好的聊天机器人;在“去束缚”增益上摘取许多明显的低垂果实,应会将我们从聊天机器人带到智能体,从工具带到看起来更像直接替代远程工人的东西。
The upshot is pretty simple. GPT-2 to GPT-4—from models that were impressive for sometimes managing to string together a few coherent sentences, to models that ace high-school exams—was not a one-time gain. We are racing through the OOMs extremely rapidly, and the numbers indicate we should expect another ~100,000x effective compute scaleup—resulting in another GPT-2-to-GPT-4-sized qualitative jump—over four years. Moreover, and critically, that doesn’t just mean a better chatbot; picking the many obvious low-hanging fruit on “unhobbling” gains should take us from chatbots to agents, from a tool to something that looks more like drop-in remote worker replacements.
虽然推理很简单,但含义是惊人的。另一次这样的飞跃很可能将我们带到 AGI,带到像博士或专家一样聪明、能作为同事与我们并肩工作的模型。也许最重要的是,如果这些 AI 系统能够自动化 AI 研究本身,那将引发强烈的反馈循环——这是本系列下一篇文章的主题。
While the inference is simple, the implication is striking. Another jump like that very well could take us to AGI, to models as smart as PhDs or experts that can work beside us as coworkers. Perhaps most importantly, if these AI systems could automate AI research itself, that would set in motion intense feedback loops—the topic of the next piece in the series.
即使是现在,几乎没有人将这一切纳入考量。但一旦你退后一步观察趋势,对 AI 的态势感知其实并不难。如果你一直对 AI 能力感到惊讶,那就开始数 OOM 吧。
Even now, barely anyone is pricing all this in. But situational awareness on AI isn’t actually that hard, once you step back and look at the trends. If you keep being surprised by AI capabilities, just start counting the OOMs.
我们现在拥有的机器,基本上可以像人类一样与之交谈。这充分证明了人类适应能力的强大,以至于这一切看起来如此平常,我们已对进步的节奏习以为常。但值得退后一步,审视一下仅仅过去几年所取得的进展。
We have machines now that we can basically talk to like humans. It’s a remarkable testament to the human capacity to adjust that this seems normal, that we’ve become inured to the pace of progress. But it’s worth stepping back and looking at the progress of just the last few years.
让我提醒你,在 GPT-4 问世前的短短约 4 年(!)里,我们取得了多大的进步。
Let me remind you of how far we came in just the ~4 (!) years leading up to GPT-4.
GPT-2(2019 年)~ 学龄前儿童:“哇,它能拼凑出几句像样的句子。”当时,它生成的一个关于安第斯山脉独角兽的半连贯故事(经过精心挑选的示例)令人印象深刻。然而,GPT-2 连数到 5 都容易出错;
GPT-2 (2019) ~ _preschooler_: “Wow, it can string together a few plausible sentences.” A very-cherry-picked example of a semi-coherent story about unicorns in the Andes it generated was incredibly impressive at the time. And yet GPT-2 could barely count to 5 without getting tripped up;
在总结文章时,它仅略优于从文章中随机选取 3 个句子。
2 when summarizing an article, it just barely outperformed selecting 3 random sentences from the article.3
当时人们对 GPT-2 印象深刻的一些例子。左图:GPT-2 在极其基础的阅读理解问题上表现尚可。右图:在一个精心挑选的样本(10 次尝试中最佳)中,GPT-2 能写出一个半连贯的段落,说一些关于内战的部分相关的内容。
_Someexamplesof what people found impressive about GPT-2 at the time. Left: GPT-2 does an ok job on extremely basic reading comprehension questions. Right: In a cherry-picked sample (best of 10 tries), GPT-2 can write a semi-coherent paragraph that says some semi-relevant things about the Civil War._
将 AI 能力与人类智能进行比较是困难且有缺陷的,但我认为考虑这里的类比是有启发性的,即使它非常不完美。GPT-2 因其语言掌控能力以及偶尔生成半连贯段落或正确回答简单事实问题的能力而令人震惊。这相当于学龄前儿童令人印象深刻的表现。
Comparing AI capabilities with human intelligence is difficult and flawed, but I think it’s informative to consider the analogy here, even if it’s highly imperfect. GPT-2 was shocking for its command of language, and its ability to occasionally generate a semi-cohesive paragraph, or occasionally answer simple factual questions correctly. It’s what would have been impressive for a preschooler.
GPT-3(2020)~ 小学生:“哇,仅凭几个少样本示例,它就能完成一些简单的有用任务。”它开始更稳定地保持多个段落的连贯性,并能纠正语法和进行非常基础的算术。首次,它在一些狭窄领域具有商业用途:例如,GPT-3 可以为 SEO 和营销生成简单的文案。
4 ~ _elementary schooler_: “Wow, with just some few-shot examples it can do some simple useful tasks.” It started being cohesive over even multiple paragraphs much more consistently, and could correct grammar and do some very basic arithmetic. For the first time, it was also commercially useful in a few narrow ways: for example, GPT-3 could generate simple copy for SEO and marketing.
当时人们对 GPT-3 印象深刻的一些例子。顶部:在简单指令后,GPT-3 能在一个新句子中使用一个生造词。左下:GPT-3 能进行丰富的故事讲述互动。右下:GPT-3 能生成一些非常简单的代码。
_Some examples of what people found impressive about GPT-3 at the time. Top: After a simple instruction, GPT-3 can use a made-up word in a new sentence. Bottom-left: GPT-3 can engage in rich storytelling back-and-forth. Bottom-right: GPT-3 can generate some very simple code._
同样,这种比较并不完美,但 GPT-3 令人印象深刻的地方或许相当于小学生:它能写一些基础诗歌,讲述更丰富连贯的故事,开始进行初步编码,能相当可靠地从简单指令和演示中学习,等等。
Again, the comparison is imperfect, but what impressed people about GPT-3 is perhaps what would have been impressive for an elementary schooler: it wrote some basic poetry, could tell richer and coherent stories, could start to do rudimentary coding, could fairly reliably learn from simple instructions and demonstrations, and so on.
GPT-4(2023 年)~ 聪明的高中生:“哇,它能编写相当复杂的代码并迭代调试,能就复杂主题进行智能而深刻的写作,能推理困难的高中竞赛数学题,在能进行的各种测试中击败绝大多数高中生,等等。”从代码到数学再到费米估算,它能思考和推理。GPT-4 现在在我的日常任务中很有用,从帮助编写代码到修改草稿。
GPT-4 (2023) ~ _smart high schooler_: “Wow, it can write pretty sophisticated code and iteratively debug, it can write intelligently and sophisticatedly about complicated subjects, it can reason through difficult high-school competition math, it’s beating the vast majority of high schoolers on whatever tests we can give it, etc.” From code to math to Fermi estimates, it can think and reason. GPT-4 is now useful in my daily tasks, from helping write code to revising drafts.
GPT-4 发布时人们印象深刻的一些例子,来自《AGI 的火花》论文。顶部:它编写非常复杂的代码(生成中间所示的图表),并能推理非平凡的数学问题。左下:解决一个 AP 数学问题。右下:解决一个相当复杂的编程问题。更多关于 GPT-4 能力探索的有趣摘录请见此处。
Some of what people found impressive about GPT-4 when it was released, from the “Sparks of AGI” paper. Top: It’s writing very complicated code (producing the plots shown in the middle) and can reason through nontrivial math problems. Bottom-left: Solving an AP math problem. Bottom-right: Solving a fairly complex coding problem. More interesting excerpts from that exploration of GPT-4’s capabilities here.
从 AP 考试到 SAT,GPT-4 的得分都优于绝大多数高中生。
On everything from AP exams to the SAT, GPT-4 scores better than the vast majority of high schoolers.
当然,即使是 GPT-4 也仍然有些不均衡;在某些任务上它比聪明的高中生好得多,而在其他任务上它尚不能完成。话虽如此,我倾向于认为这些限制大多源于模型仍被明显束缚的方式,我将在后面深入讨论。原始智能(大部分)已经存在,即使模型仍受到人为限制;需要额外的工作来解锁模型在应用中充分应用这种原始智能的能力。
Of course, even GPT-4 is still somewhat uneven; for some tasks it’s much better than smart high-schoolers, while there are other tasks it can’t yet do. That said, I tend to think most of these limitations come down to obvious ways models are still hobbled, as I’ll discuss in-depth later. The raw intelligence is (mostly) there, even if the models are still artificially constrained; it’ll take extra work to unlock models being able to fully apply that raw intelligence across applications.
仅仅四年间的进步。你在这条线上处于什么位置?
_Progress over just four years. Where are__you__on this line?_
过去十年,深度学习进步的速度简直非同寻常。仅仅十年前,深度学习系统能够识别简单图像就已经是革命性的。如今,我们不断尝试提出新颖、难度更高的测试,但每一个新基准很快就被攻克。过去,攻克广泛使用的基准需要几十年;现在感觉只需要几个月。
The pace of deep learning progress in the last decade has simply been extraordinary. A mere decade ago it was revolutionary for a deep learning system to identify simple images. Today, we keep trying to come up with novel, ever harder tests, and yet each new benchmark is quickly cracked. It used to take decades to crack widely-used benchmarks; now it feels like mere months.
_深度学习系统在许多领域正迅速达到或超越人类水平。图表:Our World in Data_
_Deep learning systems are rapidly reaching or exceeding human-level in many domains. Graphic:Our World in Data_
我们几乎快要用完基准测试了。举个例子,几年前(2020 年),我的朋友 Dan 和 Collin 创建了一个名为 MMLU 的基准测试。他们希望最终能做出一个经得起时间考验的基准,相当于我们给高中生和大学生设置的所有最难的考试。仅仅三年后,它基本上就被解决了:像 GPT-4 和 Gemini 这样的模型得分约为 90%。
We’re literally running out of benchmarks. As an anecdote, my friends Dan and Collin made a benchmark called MMLU a few years ago, in 2020. They hoped to finally make a benchmark that would stand the test of time, equivalent to all the hardest exams we give high school and college students. Just three years later, it’s basically solved: models like GPT-4 and Gemini get ~90%.
更广泛地说,GPT-4 几乎攻克了所有标准的高中和大学能力测试。
More broadly, GPT-4 mostly cracks all the standard high school and college aptitude tests.
5 (而且,从 GPT-3.5 到 GPT-4 仅一年时间,我们常常从远低于人类中位数水平跃升至人类范围的顶端。)
5 (And even the one year from GPT-3.5 to GPT-4 often took us from well below median human performance to the top of the human range.)
_GPT-4 在标准化考试中的得分。还要注意从 GPT-3.5 到 GPT-4 在这些测试的人类百分位数上的巨大跃升,通常从远低于人类中位数到人类范围的顶端。(而且这是 GPT-3.5,一个相当新的模型,发布时间比 GPT-4 早不到一年,而不是我们之前提到的笨拙的、小学水平的 GPT-3!)_
_GPT-4 scores on standardized tests. Note also the large jump from GPT-3.5 to GPT-4 in human percentile on these tests, often from well below the median human to the very top of the human range. (And this is GPT-3.5, a fairly recent model released less than a year before GPT-4, not the clunky old elementary-school-level GPT-3 we were talking about earlier!)_
_灰色:专业预测,于 2021 年 8 月做出,针对 2022 年 6 月在 MATH 基准测试(高中数学竞赛中的困难数学问题)上的表现。红星:到 2022 年 6 月的实际最先进性能,远远超过预测者给出的上限范围。中位数机器学习研究者甚至更为悲观。_
_Gray: Professional forecasts, made in August 2021, for June 2022 performance on the MATH benchmark (difficult mathematics problems from high-school math competitions). Red star: actual state-of-the-art performance by June 2022, far exceeding even the upper range forecasters gave. The median ML researcher waseven more pessimistic._
或者考虑一下 MATH 基准测试,这是一组来自高中数学竞赛的困难数学问题。
Or consider the MATH benchmark, a set of difficult mathematics problems from high-school math competitions.
6 当该基准在 2021 年发布时,最好的模型只答对了约 5%的问题。原始论文指出:“此外,我们发现,如果缩放趋势继续下去,仅仅增加预算和模型参数数量对于实现强大的数学推理是不切实际的……要在数学问题解决上取得更多进展,我们可能需要来自更广泛研究社区的新算法突破”——他们认为我们需要根本性的新突破才能解决 MATH。一项对机器学习研究者的调查预测未来几年进展甚微;7 然而,仅仅一年内(到 2022 年年中),最好的模型就从约 5%的准确率提升到了 50%;现在,MATH 基本上已被解决,最新性能超过 90%。
6 When the benchmark was released in 2021, the best models only got ~5% of problems right. And the original paper noted: “Moreover, we find that simply increasing budgets and model parameter counts will be impractical for achieving strong mathematical reasoning if scaling trends continue […]. To have more traction on mathematical problem solving we will likely need new algorithmic advancements from the broader research community”—we would need fundamental new breakthroughs to solve MATH, or so they thought. A survey of ML researchers predicted minimal progress over the coming years;7 and yet within just a year (by mid-2022), the best models went from ~5% to 50% accuracy; now, MATH is basically solved, with recent performance over 90%.
一次又一次,年复一年,怀疑论者声称“深度学习无法做到 X”,但很快就被证明是错误的。
Over and over again, year after year, skeptics have claimed “deep learning won’t be able to do X” and have been quickly proven wrong.
8_如果说我们从过去十年的人工智能中学到了一个教训,那就是永远不要押注深度学习会失败。_
8_If there’s one lesson we’ve learned from the past decade of AI, it’s that you should never bet against deep learning._
现在,最难解决的未攻克基准是像 GPQA 这样的测试,这是一组博士级别的生物学、化学和物理学问题。许多问题对我来说就像天书,即使是其他科学领域的博士,花 30 多分钟用谷歌搜索,得分也勉强高于随机概率。Claude 3 Opus 目前得分约为 60%,
Now the hardest unsolved benchmarks are tests like GPQA, a set of PhD-level biology, chemistry, and physics questions. Many of the questions read like gibberish to me, and even PhDs in other scientific fields spending 30+ minutes with Google barely score above random chance. Claude 3 Opus currently gets ~60%,
9 而领域内的博士得分约为 80%——我预计这个基准也会在接下来的一两代模型中被攻克。
9 compared to in-domain PhDs who get ~80%—and I expect this benchmark to fall as well, in the next generation or two.
_GPQA 问题示例。模型已经比我更擅长这个了,我们可能很快就会达到专家博士水平……_
_Example GPQA questions. Models are already better at this than I am, and we’ll probably crack expert-PhD-level soon…_
这是怎么发生的?深度学习的魔力在于它_就是有效_——而且趋势线惊人地一致,尽管一路上都有反对者。
How did this happen? The magic of deep learning is that it _just works_—and the trendlines have been astonishingly consistent, despite naysayers at every turn.
_以 OpenAI Sora 为例,Scaling 算力的效果。_
_The effects of scaling compute, in the example ofOpenAI Sora._
_每增加一个 OOM 的有效算力,模型都会可预测地、可靠地变得更好。_
_With each OOM of effective compute, models predictably, reliably get better._
10 如果我们能计算 OOM 的数量,我们就能(大致、定性地)推断能力的提升。11 这就是少数有远见的人预见到 GPT-4 的方式。
10 If we can count the OOMs, we can (roughly, qualitatively) extrapolate capability improvements.11 That’s how a few prescient individuals saw GPT-4 coming.
我们可以将从 GPT-2 到 GPT-4 这四年间的进展分解为三类规模扩张:
We can decompose the progress in the four years from GPT-2 to GPT-4 into three categories of scaleups:
1. _算力_:我们使用更大的计算机来训练这些模型。
1. _Compute_: We’re using much bigger computers to train these models.
2. _算法效率_:算法进步持续不断。其中许多充当“算力倍增器”,我们可以将它们统一到不断增长的_有效算力_的尺度上。
2. _Algorithmic efficiencies_: There’s a continuous trend of algorithmic progress. Many of these act as “compute multipliers,” and we can put them on a unified scale of growing _effective compute_.
3. _”解锁”增益_:默认情况下,模型学习了许多惊人的原始能力,但它们以各种愚蠢的方式受到限制,限制了其实用价值。通过简单的算法改进,如基于人类反馈的强化学习(RLHF)、思维链(CoT)、工具和脚手架,我们可以解锁显著的潜在能力。
3. _”Unhobbling” gains_: By default, models learn a lot of amazing raw capabilities, but they are hobbled in all sorts of dumb ways, limiting their practical value. With simple algorithmic improvements like reinforcement learning from human feedback (RLHF), chain-of-thought (CoT), tools, and scaffolding, we can unlock significant latent capabilities.
我们可以沿着这些轴“计算 OOM 的数量”:即,以有效算力为单位追踪每个轴的规模扩张。3 倍是 0.5 个 OOM;10 倍是 1 个 OOM;30 倍是 1.5 个 OOM;100 倍是 2 个 OOM;以此类推。我们还可以看看在 GPT-4 的基础上,从 2023 年到 2027 年应该期待什么。
We can “count the OOMs” of improvement along these axes: that is, trace the scaleup for each in units of effective compute. 3x is 0.5 OOMs; 10x is 1 OOM; 30x is 1.5 OOMs; 100x is 2 OOMs; and so on. We can also look at what we should expect on top of GPT-4, from 2023 to 2027.
我将逐一介绍,但结论很明确:我们正在快速穿越 OOM。数据墙可能存在阻力,我会讨论这一点——但总体而言,到 2027 年,在 GPT-4 的基础上,我们很可能应该期待另一次 GPT-2 到 GPT-4 规模的跃升。
I’ll go through each one-by-one, but the upshot is clear: we are rapidly racing through the OOMs. There are potential headwinds in the data wall, which I’ll address—but overall, it seems likely that we should expect another GPT-2-to-GPT-4-sized jump, on top of GPT-4, by 2027.
我先从近期进展中最常被讨论的驱动因素说起:向模型投入(大量)更多算力。
I’ll start with the most commonly-discussed driver of recent progress: throwing (a lot) more compute at models.
许多人认为这仅仅是因为摩尔定律。但即使在摩尔定律的全盛时期,其速度也相对缓慢——大约每十年提升 1-1.5 个数量级。我们看到的算力规模增长要快得多——接近摩尔定律速度的 5 倍——这主要是由于巨额投资。(过去,在单个模型上花费一百万美元是一个没人会考虑的离谱想法,而现在这只是零花钱!)
Many people assume that this is simply due to Moore’s Law. But even in the old days when Moore’s Law was in its heyday, it was comparatively _glacial_—perhaps 1-1.5 OOMs per decade. We are seeing much more rapid scaleups in compute—close to 5x the speed of Moore’s law—instead because of mammoth investment. (Spending even a million dollars on a single model used to be an outrageous thought nobody would entertain, and now that’s pocket change!)
GPT-4(2023 年)8e24 至 4e25 FLOP+ ~1.5–2 个数量级
GPT-4 (2023)8e24 to 4e25 FLOP+ ~1.5–2 OOMs
Epoch AI 对 GPT-2 到 GPT-4 算力的估算
_Estimates of compute for GPT-2 to GPT-4 byEpoch AI_
我们可以利用 Epoch AI(一个因其对 AI 趋势的出色分析而广受尊敬的来源)的公开估算来追溯 2019 年至 2023 年的算力规模增长。从 GPT-2 到 GPT-3 是一次快速的规模扩张;存在大量的算力盈余,从较小的实验扩展到使用整个数据中心来训练大型语言模型。从 GPT-3 到 GPT-4 的规模扩张中,我们过渡到了现代模式:必须为下一个模型构建一个全新的(大得多的)集群。然而,这种急剧增长仍在继续。总体而言,Epoch AI 的估算表明,GPT-4 训练使用的原始算力大约是 GPT-2 的 3000 到 10000 倍。
We can use public estimates from Epoch AI (a source widely respected for its excellent analysis of AI trends) to trace the compute scaleup from 2019 to 2023. GPT-2 to GPT-3 was a quick scaleup; there was a large overhang of compute, scaling from a smaller experiment to using an entire datacenter to train a large language model. With the scaleup from GPT-3 to GPT-4, we transitioned to the modern regime: having to build an entirely new (much bigger) cluster for the next model. And yet the dramatic growth continued. Overall, Epoch AI estimates suggest that GPT-4 training used ~3,000x-10,000x more raw compute than GPT-2.
从宏观上看,这只是一个长期趋势的延续。在过去十五年里,主要由于投资的大规模增长(以及以 GPU 和 TPU 形式为 AI 工作负载专门化的芯片),前沿 AI 系统使用的训练算力大约以每年 0.5 个数量级的速度增长。
In broad strokes, this is just the continuation of a longer-running trend. For the last decade and a half, primarily because of broad scaleups in investment (and specializing chips for AI workloads in the form of GPUs and TPUs), the training compute used for frontier AI systems has grown at roughly ~0.5 OOMs/year.
随时间推移,重要深度学习模型的训练算力。来源:Epoch AI
_Training compute of notable deep learning models over time. Source:Epoch AI_
从 GPT-2 到 GPT-3 在一年内的算力规模增长是一种不寻常的盈余,但所有迹象都表明长期趋势将持续。旧金山传闻圈中充斥着关于巨额 GPU 订单的戏剧性故事。所涉及的投资将是巨大的——但这些已经在进行中。我将在本系列的后续部分 IIIa. 竞逐万亿美元集群中进一步探讨这一点;基于该分析,到 2027 年底,再增加 2 个数量级的算力(一个数百亿美元的集群)似乎非常可能;甚至接近 3 个数量级算力(超过 1000 亿美元)的集群似乎也是可行的(并且据传微软/OpenAI 正在筹划中)。
The compute scaleup from GPT-2 to GPT-3 in a year was an unusual overhang, but all the indications are that the longer-run trend will continue. The SF-rumor-mill is abuzz with dramatic tales of huge GPU orders. The investments involved will be extraordinary—but they are in motion. I go into this more later in the series, in IIIa. Racing to the Trillion-Dollar Cluster; based on that analysis, an additional 2 OOMs of compute (a cluster in the $10s of billions) seems very likely to happen by the end of 2027; even a cluster closer to +3 OOMs of compute ($100 billion+) seems plausible (and is rumored to be in the works at Microsoft/OpenAI).
尽管对算力的大规模投资备受关注,但算法进步可能同样是进步的重要驱动力(并且一直被严重低估)。
While massive investments in compute get all the attention, algorithmic progress is probably a similarly important driver of progress (and has been dramatically underrated).
要了解算法进步能带来多大的影响,可以看看下面这个例子:在短短两年内,在 MATH 基准测试(高中数学竞赛题)上达到约 50%准确率的成本下降了。作为对比,一位不太喜欢数学的计算机科学博士生得分是 40%,所以这已经相当不错了。推理效率在不到两年内提高了近 3 个数量级——1000 倍。
To see just how big of a deal algorithmic progress can be, consider the following illustration of the drop in price to attain ~50% accuracy on the MATH benchmark (high school competition math) over just _two years._ (For comparison, a computer science PhD student who didn’t particularly like math scored 40%, so this is already quite good.)Inference efficiency improved by nearly 3 OOMs—1,000x—in less than two years.
_达到约 50% MATH 性能的相对推理成本的粗略估计。_
_Rough estimate on relative inference cost of attaining ~50% MATH performance._12
虽然这些只是推理效率的数据(可能不一定对应训练效率的提升,后者更难从公开数据中推断),但它们清楚地表明,算法进步的可能性和实际发生量都非常巨大。
Though these are numbers just for inference efficiency (which may or may not correspond to training efficiency improvements, where numbers are harder to infer from public data), they make clear there is an _enormous_ amount of algorithmic progress possible and happening.
在这篇文章中,我将区分两种算法进步。首先,我将介绍“范式内”的算法改进——这些改进只是简单地带来更好的基础模型,并直接充当_算力效率_或_算力倍增器_。例如,一个更好的算法可能让我们用 10 倍少的训练算力达到相同的性能。反过来,这相当于有效算力增加了 10 倍(1 个数量级)。(稍后,我将介绍“解绑”,你可以将其视为“范式扩展/应用扩展”的算法进步,它解锁了基础模型的能力。)
In this piece, I’ll separate out two kinds of algorithmic progress. Here, I’ll start by covering “within-paradigm” algorithmic improvements—those that simply result in better base models, and that straightforwardly act as _compute efficiencies_ or _compute multipliers_. For example, a better algorithm might allow us to achieve the same performance but with 10x less training compute. In turn, that would act as a 10x (1 OOM) increase in _effective compute_. (Later, I’ll cover “unhobbling,” which you can think of as “paradigm-expanding/application-expanding” algorithmic progress that unlocks capabilities of base models.)
如果我们退一步看长期趋势,似乎会发现新算法改进以相当稳定的速率出现。单个发现似乎是随机的,每一步都似乎有不可逾越的障碍——但长期趋势线是可预测的,在图表上是一条直线。相信趋势线。
If we step back and look at the long-term trends, we seem to find new algorithmic improvements at a fairly consistent rate. Individual discoveries seem random, and at every turn, there seem insurmountable obstacles—but the long-run trendline is predictable, a straight line on a graph. Trust the trendline.
我们有 ImageNet 的最佳数据(那里的算法研究大多是公开的,并且我们有长达十年的数据),在 2012 年至 2021 年的 9 年间,算力效率持续提高了大约每年 0.5 个数量级。
We have the best data for ImageNet (where algorithmic research has been mostly public and we have data stretching back a decade), for which we have consistently improved compute efficiency by roughly ~0.5 OOMs/year across the 9-year period between 2012 and 2021.
_我们可以衡量算法进步:与 2012 年相比,2021 年训练一个具有相同性能的模型需要多少算力?我们看到算法效率的趋势约为每年 0.5 个数量级。来源:Erdil and Besiroglu 2022。_
_We can measure algorithmic progress: how much less compute is needed in 2021 compared to 2012 to train a model with the same performance? We see a trend of ~0.5 OOMs/yr of algorithmic efficiency. Source:Erdil and Besiroglu 2022._
这意义重大:这意味着 4 年后,我们可以用大约 100 倍少的算力达到相同的性能(同时,用相同的算力可以获得高得多的性能!)。
That’s a huge deal: that means 4 years later, we can achieve the same performance for ~100x less compute (and concomitantly, much higher performance for the same compute!).
不幸的是,由于实验室不公布内部数据,衡量过去四年前沿大语言模型的算法进步更加困难。EpochAI 有新的工作将他们在 ImageNet 上的结果复制到语言建模上,并估计 2012 年至 2023 年间大语言模型的算法效率趋势约为每年 0.5 个数量级。(不过这个误差范围更大,并且没有捕捉到一些最近的收益,因为领先的实验室已经停止公布他们的算法效率。)
Unfortunately, since labs don’t publish internal data on this, it’s harder to measure algorithmic progress for frontier LLMs over the last four years. EpochAI has new work replicating their results on ImageNet for language modeling, and estimate a similar ~0.5 OOMs/year of algorithmic efficiency trend in LLMs from 2012 to 2023. (This has wider error bars though, and doesn’t capture some more recent gains, since the leading labs have stopped publishing their algorithmic efficiencies.)
_Epoch AI 对语言建模中算法效率的估计。他们的估计表明,我们在 8 年内取得了约 4 个数量级的效率提升。_
_Estimates by Epoch AI of algorithmic efficiencies in language modeling. Their estimates suggest we’ve made ~4 OOMs of efficiency gains in 8 years._
更直接地看过去四年,从 GPT-2 到 GPT-3 基本上是一次简单的规模扩展(根据论文),但自 GPT-3 以来,有许多公开已知和可推断的收益:
More directly looking at the last 4 years, GPT-2 to GPT-3 was basically a simple scaleup (according to the paper), but there have been many publicly-known and publicly-inferable gains since GPT-3:
* 我们可以从 API 成本推断收益:
* We can infer gains from API costs: 13
* GPT-4 在发布时,成本与 GPT-3 发布时大致相同,尽管性能有了巨大的提升。(如果我们根据缩放定律做一个简单粗略的估算,这表明从 GPT-3 到 GPT-4,大约一半的有效算力增长来自算法改进。)
* GPT-4, on release, cost ~the same as GPT-3 when it was released, despite the absolutely enormous performance increase. 14 (If we do a naive and oversimplified back-of-the-envelope estimate based on scaling laws, this suggests that perhaps roughly half the effective compute increase from GPT-3 to GPT-4 came from algorithmic improvements. 15)
* 自 GPT-4 发布一年以来,OpenAI 对 GPT-4 级别模型的定价随着 GPT-4o 的发布又下降了 6 倍/4 倍(输入/输出)。
* Since the GPT-4 release a year ago, OpenAI prices for GPT-4-level models have fallen another 6x/4x (input/output) with the release of GPT-4o.
* 最近发布的 Gemini 1.5 Flash 提供了介于“GPT-3.75 级别”和 GPT-4 级别之间的性能,同时成本比最初的 GPT-4 低 85 倍/57 倍(输入/输出)(非凡的收益!)。
* Gemini 1.5 Flash, recently released, offers between “GPT-3.75-level” and GPT-4-level performance, 16 while costing 85x/57x (input/output) less than the original GPT-4 (extraordinary gains!).
* Chinchilla 缩放定律带来了 3 倍以上(0.5 个数量级以上)的效率提升。
* Chinchilla scaling laws give a 3x+ (0.5 OOMs+) efficiency gain. 17
* Gemini 1.5 Pro 声称有显著的算力效率提升(性能优于 Gemini 1.0 Ultra,同时使用“显著更少”的算力),其中混合专家(MoE)是一个突出的架构变化。其他论文也声称 MoE 带来了算力的显著倍数提升。
* Gemini 1.5 Pro claimed major compute efficiency gains (outperforming Gemini 1.0 Ultra, while using “significantly less” compute), with Mixture of Experts (MoE) as a highlighted architecture change. Otherpapersalsoclaim a substantial multiple on compute from MoE.
* 在架构、数据、训练栈等方面一直有许多调整和收益。
* There have been many tweaks and gains on architecture, data, training stack, etc., all the time. 18
综合来看,公开信息表明,从 GPT-2 到 GPT-4 的跃迁包含了 1-2 个数量级的算法效率提升。
Put together, public information suggests that the GPT-2 to GPT-4 jump included 1-2 OOMs of algorithmic efficiency gains.
在 GPT-4 之后的四年里,我们应该预期这一趋势会持续:
Over the 4 years following GPT-4, we should expect the trend to continue:
平均每年 0.5 个数量级的算力效率提升,即到 2027 年相比 GPT-4 约有 2 个数量级的提升。虽然随着我们摘取低垂的果实,算力效率会变得更难找到,但 AI 实验室在寻找新算法改进上的资金和人才投入正在快速增长。(至少,公开可推断的推理成本效率似乎完全没有放缓。)在高预期下,我们甚至可能看到更根本的、类似 Transformer 的突破,带来更大的收益。
20 on average 0.5 OOMs/yr of compute efficiency, i.e. ~2 OOMs of gains compared to GPT-4 by 2027. While compute efficiencies will become harder to find as we pick the low-hanging fruit, AI lab investments in money and talent to find new algorithmic improvements are growing rapidly.21(The publicly-inferable inference cost efficiencies, at least, don’t seem to have slowed down at all.) On the high end, we could even see more fundamental, Transformer-like breakthroughs22with even bigger gains.
综合来看,这表明到 2027 年底,我们应该预期相比 GPT-4 有大约 1-3 个数量级的算法效率提升,最佳猜测可能是约 2 个数量级。
Put together, this suggests we should expect something like 1-3 OOMs of algorithmic efficiency gains (compared to GPT-4) by the end of 2027, maybe with a best guess of ~2 OOMs.
所有这一切都存在一个潜在的重要变数来源:我们正在耗尽互联网数据。这可能意味着,很快,那种简单地在更多抓取数据上预训练更大语言模型的方法可能会开始遇到严重的瓶颈。
There is a potentially important source of variance for all of this: we’re running out of internet data. That could mean that, very soon, the naive approach to pretraining larger language models on more scraped data could start hitting serious bottlenecks.
前沿模型已经在大部分互联网数据上进行了训练。例如,Llama 3 在超过 15T 个 token 上进行了训练。Common Crawl 是一个用于 LLM 训练的互联网数据转储,原始数据超过 100T 个 token,但其中大部分是垃圾信息和重复内容(例如,相对简单的去重后得到 30T 个 token,这意味着 Llama 3 已经基本上用完了所有数据)。此外,对于代码等更具体的领域,可用的 token 更少,例如,公共 GitHub 仓库估计只有几万亿个 token。
Frontier models are already trained on much of the internet. Llama 3, for example, was trained on over 15T tokens. Common Crawl, a dump of much of the internet used for LLM training, is >100T tokens raw, though much of that is spam and duplication (e.g., a relatively simple deduplication leads to 30T tokens, implying Llama 3 would already be using basically all the data). Moreover, for more specific domains like code, there are many fewer tokens still, e.g. public github repos are estimated to be in low trillions of tokens.
通过重复数据可以在一定程度上进一步扩展,但学术研究表明,重复的效果有限,发现在 16 个 epoch(16 倍重复)之后,收益会极快地降至零。在某个点上,即使有更多的(有效)算力,由于数据约束,让模型变得更好也会变得更加困难。这一点不容低估:我们一直乘着 Scaling(规模扩张)曲线,乘着语言模型预训练范式的浪潮,如果没有新的东西出现,这个范式(至少天真地看)将会耗尽。尽管有巨额投资,我们也会陷入平台期。据传所有实验室都在进行大规模的研究押注,寻找新的算法改进或方法来绕过这一障碍。研究人员据称正在尝试多种策略,从合成数据到自我对弈和强化学习方法。业内人士似乎非常乐观:Anthropic 的 CEO Dario Amodei 最近在一个播客中说:“如果非常天真地看,我们离数据耗尽并不远……我猜测这不会成为障碍……有很多不同的方法可以做到这一点。”当然,任何这方面的研究成果都是专有的,如今不再公开发表。
You can go somewhat further by repeating data, but academic work on this suggests that repetition only gets you so far, finding that after 16 epochs (a 16-fold repetition), returns diminish extremely fast to nil. At some point, even with more (effective) compute, making your models better can become much tougher because of the data constraint. This isn’t to be understated: we’ve been riding the scaling curves, riding the wave of the language-modeling-pretraining-paradigm, and without something new here, this paradigm will (at least naively) run out. Despite the massive investments, we’d plateau. All of the labs are rumored to be making massive research bets on new algorithmic improvements or approaches to get around this. Researchers are purportedly trying many strategies, from synthetic data to self-play and RL approaches. Industry insiders seem to be very bullish: Dario Amodei (CEO of Anthropic) recently said on a podcast: “if you look at it very naively we’re not that far from running out of data […] My guess is that this will not be a blocker […] There’s just many different ways to do it.” Of course, any research results on this are proprietary and not being published these days.
除了业内人士的乐观态度,我认为还有一个强有力的直观理由说明为什么应该有可能找到训练模型的方法,使其具有更好的样本效率(算法改进,让它们从有限数据中学到更多)。想想你或我如何从一本非常密集的数学教科书中学习:
In addition to insider bullishness, I think there’s a strong intuitive case for why it should be possible to find ways to train models with much better sample efficiency (algorithmic improvements that let them learn more from limited data). Consider how you or I would learn from a really dense math textbook:
* 现代 LLM 在训练期间所做的,本质上是非常快速地浏览教科书,文字飞速掠过,没有花费太多脑力。
* What a modern LLM does during training is, essentially, very very quickly skim the textbook, the words just _flying by_, not spending much brain power on it.
* 相反,当你或我阅读那本数学教科书时,我们慢慢地读几页;然后在脑海中就材料进行内心独白,并与几个学习伙伴讨论;再读一两页;然后尝试一些练习题,失败,用不同的方法再试,获得关于这些问题的反馈,再试直到做对一道题;如此反复,直到材料最终“融会贯通”。
* Rather, when you or I read that math textbook, we read a couple pages slowly; then have an internal monologue about the material in our heads and talk about it with a few study-buddies; read another page or two; then try some practice problems, fail, try them again in a different way, get some feedback on those problems, try again until we get a problem right; and so on, until eventually the material “clicks.”
* 你或我如果只能像 LLM 那样快速浏览一本密集的数学教科书,也不会学到多少东西。
* You or I also wouldn’t learn much at all from a pass through a dense math textbook if all we could do was breeze through it like LLMs. 23
* 但也许,有办法融入人类消化一本密集数学教科书的方式,让模型从有限数据中学到更多。简而言之,这类事情——对材料进行内心独白,与学习伙伴讨论,尝试和失败直到弄懂——正是许多合成数据/自我对弈/强化学习方法试图做到的。
* But perhaps, then, there are ways to incorporate aspects of how humans would digest a dense math textbook to let the models learn much more from limited data. In a simplified sense, this sort of thing—having an internal monologue about material, having a discussion with a study-buddy, trying and failing at problems until it clicks—is what many synthetic data/self-play/RL approaches are trying to do. 24
旧的模型训练技术简单而天真,但它有效,所以没有人真正努力去攻克这些提高样本效率的方法。现在这可能会成为一个更大的约束,我们应该期待所有实验室投入数十亿美元和他们最聪明的头脑来攻克它。深度学习中的一个常见模式是,需要付出大量努力(以及许多失败的项目)才能把细节做好,但最终某个显而易见且简单的东西就是能行得通。鉴于深度学习在过去十年中成功突破了每一个所谓的障碍,我的基本判断是这次也会类似。
The old state of the art of training models was simple and naive, but it worked, so nobody really tried hard to crack these approaches to sample efficiency. Now that it may become more of a constraint, we should expect all the labs to invest billions of dollars and their smartest minds into cracking it. A common pattern in deep learning is that it takes _a lot_ of effort (and many failed projects) to get the details right, but eventually some version of the obvious and simple thing just works. Given how deep learning has managed to crash through every supposed wall over the last decade, my base case is that it will be similar here.
此外,实际上有可能攻克这些算法赌注之一(如合成数据)会显著改进模型。这里有一个直觉泵。当前的前沿模型如 Llama 3 是在互联网上训练的——而互联网大部分是垃圾,比如电子商务或 SEO 之类。许多 LLM 将其绝大部分训练算力花在这些垃圾上,而不是真正高质量的数据上(例如,人们解决困难科学问题的推理链)。想象一下,如果你能把 GPT-4 级别的算力完全花在极其高质量的数据上——那可能会是一个能力强大得多的模型。
Moreover, it actually seems possible that cracking one of these algorithmic bets like synthetic data could dramatically _improve_ models. Here’s an intuition pump. Current frontier models like Llama 3 are trained on the internet—and the internet is mostly crap, like e-commerce or SEO or whatever. Many LLMs spend the vast majority of their training compute on this crap, rather than on really high-quality data (e.g. reasoning chains of people working through difficult science problems). Imagine if you could spend GPT-4-level compute on entirely extremely high-quality data—it could be a much, much more capable model.
回顾 AlphaGo——第一个在围棋上击败世界冠军的人工智能系统,比人们认为可能的时间早了数十年——这里也很有用。
A look back at AlphaGo—the first AI system that beat the world champions at the game of Go, decades before it was thought possible—is useful here as well.
* 第一步,AlphaGo 通过模仿学习在人类专家围棋棋谱上训练。这给了它一个基础。
* In step 1, AlphaGo was trained by imitation learning on expert human Go games. This gave it a foundation.
* 第二步,AlphaGo 与自己下了数百万盘棋。这使它成为围棋超人:还记得与李世石对局中著名的第 37 手吗?那是一个极其不寻常但 brilliant 的着法,人类永远不会下出。
* In step 2, AlphaGo played millions of games against itself. This let it become superhuman at Go: remember the famous move 37 in the game against Lee Sedol, an extremely unusual but brilliant move a human would never have played.
为 LLM 开发相当于第二步的方法是克服数据墙的关键研究问题(而且,最终将是超越人类智能的关键)。
Developing the equivalent of step 2 for LLMs is a key research problem for overcoming the data wall (and, moreover, will ultimately be the key to surpassing human-level intelligence).
所有这一切都表明,数据约束似乎给预测未来几年的人工智能进展带来了很大的误差范围。事情有非常真实的可能性会停滞(LLM 可能仍然像互联网一样重要,但我们不会达到真正疯狂的 AGI)。但我认为有理由猜测实验室会攻克它,并且这样做不仅会保持 Scaling(规模扩张)曲线继续,还可能带来模型能力的巨大提升。
All of this is to say that data constraints seem to inject large error bars either way into forecasting the coming years of AI progress. There’s a very real chance things stall out (LLMs might still be as big of a deal as the internet, but we wouldn’t get to truly crazy AGI). But I think it’s reasonable to guess that the labs will crack it, and that doing so will not just keep the scaling curves going, but possibly enable huge gains in model capability.
顺便说一句,这也意味着我们应该预期未来几年不同实验室之间的差异会比今天更大。直到最近,最先进的技术都是公开的,所以每个人基本上都在做同样的事情。(新进入者或开源项目可以轻松与前沿竞争,因为配方是公开的。)现在,关键的算法思想变得越来越专有。我预计实验室的方法会更加分化,有些实验室会比其他人进步更快——即使是现在处于前沿的实验室也可能在数据墙上停滞不前,而其他实验室取得突破从而领先。开源将更难竞争。这肯定会让事情变得有趣。(而且,如果某个实验室解决了这个问题,他们的突破将是 AGI 的关键,是超级智能的关键——美国最宝贵的机密之一。)
As an aside, this also means that we should expect more variance between the different labs in coming years compared to today. Up until recently, the state of the art techniques were published, so everyone was basically doing the same thing. (And new upstarts or open source projects could easily compete with the frontier, since the recipe was published.) Now, key algorithmic ideas are becoming increasingly proprietary. I’d expect labs’ approaches to diverge much more, and some to make faster progress than others—even a lab that seems on the frontier now could get stuck on the data wall while others make a breakthrough that lets them race ahead. And open source will have a much harder time competing. It will certainly make things interesting. (And if and when a lab figures it out, their breakthrough will be the key to AGI, key to superintelligence—one of the United States’ most prized secrets.)
最后,最难量化但同样重要的改进类别:我称之为“去束缚”。
Finally, the hardest to quantify—but no less important—category of improvements: what I’ll call “unhobbling.”
想象一下,如果要求你解决一个困难的数学问题,你必须立即回答脑海中闪现的第一个念头。显然,除了最简单的问题外,你会很难应对。但直到最近,我们就是这样让大语言模型解决数学问题的。相反,我们大多数人会在草稿纸上逐步推导,从而能够解决更困难的问题。“思维链”提示为大语言模型解锁了这种能力。尽管它们拥有出色的原始能力,但在数学方面却远不如预期,因为它们以一种明显的方式被束缚了,而一个小的算法调整就能释放出更大的能力。
Imagine if when asked to solve a hard math problem, you had to instantly answer with the very first thing that came to mind. It seems obvious that you would have a hard time, except for the simplest problems. But until recently, that’s how we had LLMs solve math problems. Instead, most of us work through the problem step-by-step on a scratchpad, and are able to solve much more difficult problems that way. “Chain-of-thought” prompting unlocked that for LLMs. Despite excellent raw capabilities, they were much worse at math than they could be because they were hobbled in an obvious way, and it took a small algorithmic tweak to unlock much greater capabilities.
在过去几年中,我们在“去束缚”模型方面取得了巨大进展。这些算法改进超越了仅仅训练更好的基础模型——而且通常只使用预训练算力的一小部分——却能释放模型的能力:
We’ve made huge strides in “unhobbling” models over the past few years. These are algorithmic improvements beyond just training better base models—and often only use a fraction of pretraining compute—that unleash model capabilities:
* _基于人类反馈的强化学习(RLHF)_。基础模型拥有令人难以置信的_潜在_能力,但它们原始且极难使用。虽然普遍认为 RLHF 仅仅是审查脏话,但 RLHF 一直是使模型真正有用和具有商业价值的关键(而不是让模型预测随机的互联网文本,而是让它们实际应用能力来回答你的问题!)。这就是 ChatGPT 的魔力——精心设计的 RLHF 首次使模型对真实用户可用且有用。原始的 InstructGPT 论文对此有很好的量化:在人类评分者偏好方面,一个经过 RLHF 的小模型相当于一个未经 RLHF 的大 100 倍以上的模型。
* _Reinforcement learning from human feedback (RLHF)_. Base models have incredible _latent_ capabilities, 26but they’re raw and incredibly hard to work with. While the popular conception of RLHF is that it merely censors swear words, RLHF has been key to making models actually useful and commercially valuable (rather than making models predict random internet text, get them to actually apply their capabilities to try to answer your question!). This was the magic of ChatGPT—well-done RLHF made models usable and useful to real people for the first time. The original InstructGPT paper has a great quantification of this: an RLHF’d small model was equivalent to a non-RLHF’d >100x larger model in terms of human rater preference.
* _思维链_(CoT)。如前所述。CoT 在两年前才开始广泛使用,在数学/推理问题上可以提供相当于有效算力增加 10 倍以上的效果。
* _Chain of Thought_ (CoT). As discussed. CoT started being widely used just 2 years ago and can provide the equivalent of a >10x effective compute increase on math/reasoning problems.
* _脚手架_。可以看作是 CoT++:不是仅仅让模型解决问题,而是让一个模型制定攻击计划,另一个提出一系列可能的解决方案,另一个进行批评,等等。例如,在 HumanEval(编程问题)上,简单的脚手架使 GPT-3.5 能够超越未经脚手架的 GPT-4。在 SWE-Bench(解决真实世界软件工程任务的基准)上,GPT-4 只能正确解决约 2%的问题,而使用 Devin 的智能体脚手架后,这一比例跃升至 14-23%。(不过,释放智能体能力仍处于初期阶段,我稍后会进一步讨论。)
* _Scaffolding_. Think of CoT++: rather than just asking a model to solve a problem, have one model make a plan of attack, have another propose a bunch of possible solutions, have another critique it, and so on. For example, on HumanEval (coding problems), simple scaffolding enables GPT-3.5 to outperform un-scaffolded GPT-4. On SWE-Bench (a benchmark of solving real-world software engineering tasks), GPT-4 can only solve ~2% correctly, while with Devin’s agent scaffolding it jumps to 14-23%. (Unlocking agency is only in its infancy though, as I’ll discuss more later.)
* _工具_:想象一下,如果人类不允许使用计算器或电脑。我们在这方面才刚刚起步,但 ChatGPT 现在可以使用网页浏览器、运行一些代码等。
* _Tools:_ Imagine if humans weren’t allowed to use calculators or computers. We’re only at the beginning here, but ChatGPT can now use a web browser, run some code, and so on.
* _上下文长度_。模型从 2k token 上下文(GPT-3)发展到 32k 上下文(GPT-4 发布时),再到 100 万+上下文(Gemini 1.5 Pro)。这是一个_巨大的_进步。一个较小的基础模型,如果拥有例如 10 万 token 的相关上下文,可以超越一个更大但只有例如 4k 相关 token 上下文的模型——更多的上下文实际上是一种巨大的算力效率提升。更一般地说,上下文是解锁这些模型许多应用的关键:例如,许多编程应用需要理解代码库的大部分内容才能有效地贡献新代码;或者,如果你使用模型帮助你在工作中撰写文档,它确实需要来自大量相关内部文档和对话的上下文。Gemini 1.5 Pro 拥有 100 万+ token 的上下文,甚至能够仅通过将词典和语法参考资料放入上下文,从零开始学习一门新语言(一种互联网上不存在的低资源语言)!
* _Context length_. Models have gone from 2k token context (GPT-3) to 32k context (GPT-4 release) to 1M+ context (Gemini 1.5 Pro). This is a _huge_ deal. A much smaller base model with, say, 100k tokens of relevant context can outperform a model that is much larger but only has, say, 4k relevant tokens of context—more context is effectively a large compute efficiency gain. 27 More generally, context is key to unlocking many applications of these models: for example, many coding applications require understanding large parts of a codebase in order to usefully contribute new code; or, if you’re using a model to help you write a document at work, it really needs the context from lots of related internal docs and conversations. Gemini 1.5 Pro, with its 1M+ token context, was even able to learn a new language (a low-resource language not on the internet) from scratch, just by putting a dictionary and grammar reference materials in context!
* _后训练改进_。根据 John Schulman 的说法,当前的 GPT-4 与最初发布时的 GPT-4 相比有了显著改进,这是由于后训练改进释放了潜在的模型能力:在推理评估上取得了实质性进展(例如,MATH 从约 50%提高到 72%,GPQA 从约 40%提高到约 50%),在 LMSys 排行榜上,Elo 分数提高了近 100 分(相当于 Claude 3 Haiku 和更大的 Claude 3 Opus 之间的 Elo 差异,这两个模型的价格相差约 50 倍)。
* _Posttraining improvements._ The current GPT-4 has substantially improved compared to the original GPT-4 when released, according to John Schulman due to posttraining improvements that unlocked latent model capability: on reasoning evals it’s made substantial gains (e.g., ~50% -> 72% on MATH, ~40% to ~50% on GPQA) and on the LMSys leaderboard, it’s made nearly 100-point elo jump (comparable to the difference in elo between Claude 3 Haiku and the much larger Claude 3 Opus, models that have a ~50x price difference).
Epoch AI 对其中一些技术(如脚手架、工具使用等)的调查发现,这类技术通常可以在许多基准上实现 5-30 倍的有效算力增益。METR(一个评估模型的组织)同样发现,在相同 GPT-4 基础模型上通过去束缚,其一组智能体任务的性能得到了非常大的提升:从仅基础模型的 5%,到发布时经过后训练的 GPT-4 的 20%,再到今天通过更好的后训练、工具和智能体脚手架达到近 40%。
A survey by Epoch AI of some of these techniques, like scaffolding, tool use, and so on, finds that techniques like this can typically result in effective compute gains of 5-30x on many benchmarks. METR (an organization that evaluates models) similarly found very large performance improvements on their set of agentic tasks, via unhobbling from the same GPT-4 base model: from 5% with just the base model, to 20% with the GPT-4 as posttrained on release, to nearly 40% today from better posttraining, tools, and agent scaffolding.
_METR 智能体任务性能随时间通过更好的去束缚而提升。来源:模型评估与威胁研究_
_Performance on METR’s agentic tasks, over time via better unhobbling. Source:Model Evaluation and Threat Research_
虽然很难将这些与算力和算法效率统一到有效的算力尺度上,但很明显这些是_巨大的_收益,至少与算力扩展和算法效率大致相当。(这也凸显了算法进步的核心作用:每年 0.5 个数量级的算力效率已经显著,但只是故事的一部分,与去束缚结合起来,算法进步总体上可能占据了当前趋势收益的大部分。)
While it’s hard to put these on a unified effective compute scale with compute and algorithmic efficiencies, it’s clear these are _huge_ gains, at least on a roughly similar magnitude as the compute scaleup and algorithmic efficiencies. (It also highlights the central role of algorithmic progress: the 0.5 OOMs/year of compute efficiencies, already significant, are only part of the story, and put together with unhobbling algorithmic progress overall is maybe even a majority of the gains on the current trend.)
“去束缚”实际上使这些模型变得有用——我认为,今天许多商业应用受到阻碍的原因正是需要进一步的“去束缚”。事实上,_今天的模型仍然被严重束缚_!例如:
“Unhobbling” is what actually enabled these models to become useful—and I’d argue that much of what is holding back many commercial applications today is the need for further “unhobbling” of this sort. Indeed, _models today are still incredibly hobbled_!For example:
* 它们不能使用电脑(仍然只有非常有限的工具)。
* They can’t use a computer (they still only have very limited tools).
* 它们大多仍然在说话前不思考。当你要求 ChatGPT 写一篇文章时,这就像期望人类通过初始的意识流来写文章。
* They still mostly don’t think before they speak. When you ask ChatGPT to write an essay, that’s like expecting a human to write an essay via their initial stream-of-consciousness. 28
* 它们(大多)只能进行简短的来回对话,而不能离开一天或一周,思考一个问题,研究不同的方法,咨询其他人,然后给你写一份更长的报告或拉取请求。
* They can (mostly) only engage in short back-and-forth dialogues, rather than going away for a day or a week, thinking about a problem, researching different approaches, consulting other humans, and then writing you a longer report or pull request.
* 它们大多没有针对你或你的应用进行个性化(只是一个带有简短提示的通用聊天机器人,而不是拥有你公司和工作的所有相关背景)。
* They’re mostly not personalized to you or your application (just a generic chatbot with a short prompt, rather than having all the relevant background on your company and your work).
这里的可能性是巨大的,我们正在迅速摘取低垂的果实。这一点至关重要:_仅仅想象“GPT-6 ChatGPT”是完全错误的。_ 随着去束缚的持续进展,与 GPT-6 加 RLHF 相比,改进将是阶跃式的。到 2027 年,你将拥有的不是一个聊天机器人,而是一个更像智能体、像同事的东西。
The possibilities here are enormous, and we’re rapidly picking low-hanging fruit here. This is critical: _it’s completely wrong to just imagine “GPT-6 ChatGPT.”_ With continued unhobbling progress, the improvements will be step-changes compared to GPT-6 + RLHF. By 2027, rather than a chatbot, you’re going to have something that looks more like an agent, like a coworker.
未来几年雄心勃勃的“去束缚”会是什么样子?在我看来,有三个关键要素:
What could ambitious unhobbling over the coming years look like? The way I think about it, there are three key ingredients:
GPT-4 拥有足够的原始智能来完成许多人工作中的相当一部分,但它就像一个刚出现五分钟的聪明新员工:没有任何相关背景,没有阅读公司文档或 Slack 历史记录,没有与团队成员交谈过,也没有花时间了解公司内部的代码库。一个聪明的新员工在到达五分钟后并没有太大用处——但一个月后就会相当有用!似乎应该有可能,例如通过超长上下文,像我们对待新的人类同事那样“入职”模型。仅此一项就将是一个巨大的解锁。
GPT-4 has the raw smarts to do a decent chunk of many people’s jobs, but it’s sort of like a smart new hire that just showed up 5 minutes ago: it doesn’t have any relevant context, hasn’t read the company docs or Slack history or had conversations with members of the team, or spent any time understanding the company-internal codebase. A smart new hire isn’t that useful 5 minutes after arriving—but they are quite useful a month in! It seems like it should be possible, for example via very-long-context, to “onboard” models like we would a new human coworker. This alone would be a huge unlock.
2. 测试时算力过剩(推理/错误修正/系统二,用于更长周期的问题)
2.The test-time compute overhang (reasoning/error correction/system II for longer-horizon problems)
目前,模型基本上只能完成短任务:你问它们一个问题,它们给出一个答案。但这极其有限。人类完成的大多数有用的认知工作周期更长——不仅仅需要五分钟,而是数小时、数天、数周或数月。
Right now, models can basically only do short tasks: you ask them a question, and they give you an answer. But that’s extremely limiting. Most useful cognitive work humans do is longer horizon—it doesn’t just take 5 minutes, but hours, days, weeks, or months.
一个只能思考一个难题五分钟的科学家不可能取得任何科学突破。一个软件工程师如果只能在被要求时编写单个函数的骨架代码,那将非常没用——软件工程师会被赋予一个更大的任务,然后他们制定计划,理解代码库或技术工具的相关部分,编写不同的模块并逐步测试,调试错误,搜索可能的解决方案空间,最终提交一个大型拉取请求,这是数周工作的结晶。等等。
A scientist that could only think about a difficult problem for 5 minutes couldn’t make any scientific breakthroughs. A software engineer that could only write skeleton code for a single function when asked wouldn’t be very useful—software engineers are given a larger task, and they then go make a plan, understand relevant parts of the codebase or technical tools, write different modules and test them incrementally, debug errors, search over the space of possible solutions, and eventually submit a large pull request that’s the culmination of weeks of work. And so on.
本质上,存在一个巨大的_测试时算力过剩_。将每个 GPT-4 词元视为你在思考问题时内心独白的一个词。每个 GPT-4 词元都非常聪明,但目前它只能有效地使用大约数百个词元来进行连贯的思维链(就好像你只能花几分钟的内心独白/思考在一个问题或项目上)。
In essence, there is a large _test-time compute overhang._ Think of each GPT-4 token as a word of internal monologue when you think about a problem. Each GPT-4 token is quite smart, but it can currently only really effectively use on the order of ~hundreds of tokens for chains of thought coherently (effectively as though you could only spend a few minutes of internal monologue/thinking on a problem or project).
如果它能使用数百万个词元来思考和解决真正困难的问题或更大的项目呢?
What if it could use millions of tokens to think about and work on really hard problems or bigger projects?
词元数量 相当于我花在某件事上的时间……
Number of tokens Equivalent to me working on something for…
数千 半小时以上 +1 个数量级的测试时算力
1000s Half an hour+1 OOMs test-time compute
_假设人类以约 100 词元/分钟的速度思考,每周工作 40 小时,将模型“思考”的词元数量转换为人类在给定问题/项目上的时间。_
_Assuming a human thinking at ~100 tokens/minute and working 40 hours/week, translating “how long a model thinks” in tokens to human-time on a given problem/project._
即使“每词元”的智能相同,这也相当于一个聪明人在一个问题上花费_几分钟_与_几个月_的差别。我不知道你怎么想,但我在几个月内能做的事情比几分钟内多得多、多得多、多得多。如果我们能解锁模型“能够思考和解决相当于数月的工作,而不是几分钟的工作”,那将解锁_疯狂的_能力跃升。这里存在巨大的过剩,许多个数量级。
Even if the “per-token” intelligence were the same, it’d be the difference between a smart person spending a _few minutes_ vs. a _few months_ on a problem. I don’t know about you, but there’s much, much, much more I am capable of in a few months vs. a few minutes. If we could unlock “being able to think and work on something for months-equivalent, rather than a few-minutes-equivalent” for models, it would unlock an _insane_ jump in capability. There’s a huge overhang here, many OOMs worth.
目前,模型还无法做到这一点。即使最近在长上下文方面取得了进展,这种较长的上下文主要也只适用于词元的消费,而不是词元的产生——过了一段时间,模型就会偏离轨道或卡住。它还不能独自离开一段时间去解决一个问题或项目。
Right now, models can’t do this yet. Even with recent advances in long-context, this longer context mostly only works for the consumption of tokens, not the production of tokens—after a while, the model goes off the rails or gets stuck. It’s not yet able to go away for a while to work on a problem or project on its own.
但解锁测试时算力可能仅仅需要相对较小的“去束缚”算法胜利。也许少量的强化学习可以帮助模型学会错误修正(“嗯,那看起来不对,让我再检查一下”)、制定计划、搜索可能的解决方案等等。从某种意义上说,模型已经具备了大部分原始能力,它只需要额外学习一些技能来将它们整合在一起。
But unlocking test-time compute might merely be a matter of relatively small “unhobbling” algorithmic wins. Perhaps a small amount of RL helps a model learn to error correct (“hm, that doesn’t look right, let me double check that”), make plans, search over possible solutions, and so on. In a sense, the model already has most of the raw capabilities, it just needs to learn a few extra skills on top to put it all together.
本质上,我们只需要教会模型一种系统二的外循环
In essence, we just need to teach the model a sort of System II outer loop
,让它能够推理困难、长周期的项目。
30 that lets it reason through difficult, long-horizon projects.
如果我们成功教会了这个外循环,那么想象一下,不再是几段简短的聊天机器人回答,而是数百万词的流(以你快于阅读的速度出现),模型在思考问题、使用工具、尝试不同方法、进行研究、修改工作、与他人协调以及独立完成大型项目。
If we succeed at teaching this outer loop, instead of a short chatbot answer of a couple paragraphs, imagine a stream of millions of words (coming in more quickly than you can read them) as the model thinks through problems, uses tools, tries different approaches, does research, revises its work, coordinates with others, and completes big projects on its own.
在其他领域,比如棋盘游戏的 AI 系统中,已经证明可以使用更多的测试时算力(也称为推理时算力)来替代训练算力。
In other domains, like AI systems for board games, it’s been demonstrated that you can use more test-time compute (also called inference-time compute) to substitute for training compute.
Jones (2021):在 Hex 游戏中,如果给较小的模型更多的测试时算力(“更多的思考时间”),它可以表现得和更大的模型一样好。在这个领域中,他们发现,花费约 1.2 个数量级更多的测试时算力,可以获得与训练算力增加约 1 个数量级相当的模型性能。
Jones (2021): A smaller model can do as well as a much larger model at the game of Hex if you give it more test-time compute (“more time to think”). In this domain, they find that one can spend ~1.2 OOMs more compute at test-time to get performance equivalent to a model with ~1 OOMs more training compute.
如果类似的关系在我们的案例中成立,那么如果我们能解锁+4 个数量级的测试时算力,可能相当于+3 个数量级的预训练算力,即大致相当于从 GPT-3 到 GPT-4 的跃迁。(也就是说,解决这个“去束缚”问题相当于巨大的数量级扩展。)
If a similar relationship held in our case, if we could unlock + 4 OOMs of test-time compute, that might be equivalent to + 3OOMs of pretraining compute, i.e. very roughly something like the jump between GPT-3 and GPT-4. (I.e., solving this “unhobbling” would be equivalent to a huge OOM scaleup.)
这也许是三者中最直接的一个。ChatGPT 现在基本上就像一个坐在与世隔绝的盒子里、你可以发短信给它的人。虽然早期的去束缚改进教会模型使用单个孤立的工具,但我预计,随着多模态模型的出现,我们很快就能一举实现这一点:我们将简单地让模型像人类一样使用计算机。
This is perhaps the most straightforward of the three. ChatGPT right now is basically like a human that sits in an isolated box that you can text. While early unhobbling improvements teach models to use individual isolated tools, I expect that with multimodal models we will soon be able to do this in one fell swoop: we will simply enable models to use a computer like a human would.
这意味着加入你的 Zoom 会议、在线研究事物、给人发消息和发邮件、阅读共享文档、使用你的应用程序和开发工具等等。(当然,为了让模型在更长周期的循环中充分利用这一点,这需要与解锁测试时算力相辅相成。)
That means joining your Zoom calls, researching things online, messaging and emailing people, reading shared docs, using your apps and dev tooling, and so on. (Of course, for models to make the most use of this in longer-horizon loops, this will go hand-in-hand with unlocking test-time compute.)
到这一步结束时,我预计我们会得到一种看起来很像“即插即用的远程工作者”的东西。一个加入你公司、像新员工一样入职、在 Slack 上给你和同事发消息、使用你的软件、提交拉取请求的智能体,并且,在给定大型项目时,它可以像人类一样离开数周独立完成项目。你可能需要比 GPT-4 更好的基础模型来解锁这一点,但可能甚至不需要好那么多——很多潜力在于修复模型仍然被束缚的明显和基本的方式。
By the end of this, I expect us to get something that looks a lot like a _drop-in remote worker_. An agent that joins your company, is onboarded like a new human hire, messages you and colleagues on Slack and uses your softwares, makes pull requests, and that, given big projects, can do the model-equivalent of a human going away for weeks to independently complete the project. You’ll probably need somewhat better base models than GPT-4 to unlock this, but possibly not even _that_ much better—a lot of juice is in fixing the clear and basic ways models are still hobbled.
_一个非常早期的预览是 Devin,这是一个早期原型,它在通往创建全自动软件工程师的道路上解锁了模型的“智能体过剩”/“测试时算力过剩”。我不知道 Devin 在实践中表现如何,这个演示与从聊天机器人到智能体的适当去束缚所能带来的相比仍然非常有限,但它是一个有用的预告,展示了即将到来的这类事物。_
_A very early peek at what this might look like isDevin, an early prototype of unlocking the “agency-overhang”/”test-time compute overhang” on models on the path to creating a fully automated software engineer. I don’t know how well Devin works in practice, and this demo is still very limited compared to what proper chatbot → agent unhobbling would yield, but it’s a useful teaser of the sort of thing coming soon._
顺便说一句,我预计去束缚的核心地位将在商业应用方面导致一种相当有趣的“音爆”效应。从现在到即插即用的远程工作者之间的中间模型将需要大量的繁琐工作来改变工作流程和构建基础设施,以便整合并从中获得经济价值。而即插即用的远程工作者将更容易整合——只需直接让他们加入,就能自动化所有可以远程完成的工作。似乎合理的是,繁琐工作可能比去束缚花费更长时间,也就是说,等到即插即用的远程工作者能够自动化大量工作时,中间模型可能还没有被充分利用和整合——因此产生的经济价值跃升可能有些非连续。
By the way, I expect the centrality of unhobbling to lead to a somewhat interesting “sonic boom” effect in terms of commercial applications. Intermediate models between now and the drop-in remote worker will require tons of schlep to change workflows and build infrastructure to integrate and derive economic value from. The drop-in remote worker will be dramatically easier to integrate—just, well, drop them in to automate all the jobs that could be done remotely. It seems plausible that the schlep will take longer than the unhobbling, that is, by the time the drop-in remote worker is able to automate a large number of jobs, intermediate models won’t yet have been fully harnessed and integrated—so the jump in economic value generated could be somewhat discontinuous.
_对 GPT-4 前四年进展驱动因素的估计总结,以及我们对 GPT-4 后四年应有何种预期。_
_Summary of the estimates on drivers of progress in the four years preceding GPT-4, and what we should expect in the four years following GPT-4._
综合这些数字,我们大致可以预期,在 GPT-4 之后的四年内,将再次出现类似 GPT-2 到 GPT-4 规模的跃升,到 2027 年底。
Putting the numbers together, we should (roughly) expect another GPT-2-to-GPT-4-sized jump in the 4 years following GPT-4, by the end of 2027.
* GPT-2 到 GPT-4 大致是 4.5–6 个数量级的基础有效算力扩展(物理算力和算法效率),加上主要的“解禁”收益(从基础模型到聊天机器人)。
* GPT-2 to GPT-4 was roughly a 4.5–6 OOM base effective compute scaleup (physical compute and algorithmic efficiencies), plus major “unhobbling” gains (from base model to chatbot).
* 在接下来的四年里,我们应预期 3–6 个数量级的基础有效算力扩展(物理算力和算法效率)——最佳猜测可能是约 5 个数量级——加上由“解禁”带来的效用和应用上的阶跃变化(从聊天机器人到智能体/即插即用的远程工作者)。
* In the subsequent 4 years, we should expect 3–6 OOMs of base effective compute scaleup (physical compute and algorithmic efficiencies)—with perhaps a best guess of ~5 OOMs—plus step-changes in utility and applications unlocked by “unhobbling” (from chatbot to agent/drop-in remote worker).
为了更直观地理解,假设 GPT-4 训练耗时 3 个月。_到 2027 年,一家领先的 AI 实验室将能在一分钟内训练出一个 GPT-4 级别的模型。_
To put this in perspective, suppose GPT-4 training took 3 months. _In 2027, a leading AI lab will be able to train a GPT-4-level model in a minute._
31 有效算力的数量级扩展将是巨大的。
31The OOM effective compute scaleup will be dramatic.
GPT-2 到 GPT-4 让我们从大约学龄前儿童水平进步到聪明的高中生水平;从几乎无法输出几个连贯句子到通过高中考试并成为有用的编程助手。这是一个疯狂的飞跃。如果_这_是我们将再次跨越的智能差距,那会带我们走向何方?
GPT-2 to GPT-4 took us from ~preschooler to ~smart high-schooler; from barely being able to output a few cohesive sentences to acing high-school exams and being a useful coding assistant. That was an insane jump. If _this_ is the intelligence gap we’ll cover once more, where will that take us?
32 如果这把我们带向非常非常远的地方,我们不应感到惊讶。很可能,它将带我们走向能超越博士和领域内最优秀专家的模型。
32 We should not be surprised if that takes us very, very far. Likely, it will take us to models that can outperform PhDs and the best experts in a field.
(一个简洁的思考方式是:当前 AI 进展的速度大约是儿童发展速度的 3 倍。你那个 3 倍速的孩子刚刚高中毕业;它很快就会在你意识到之前抢走你的工作!)
(One neat way to think about this is that the current trend of AI progress is proceeding at roughly 3x the pace of child development. Your 3x-speed-child just graduated high school; it’ll be taking your job before you know it!)
再次强调,关键的是,不要仅仅想象一个极其聪明的 ChatGPT:解禁收益应该意味着这更像一个即插即用的远程工作者,一个极其聪明的智能体,能够推理、规划、纠错,了解你和你的公司的一切,并能独立工作数周。
Again, critically, don’t just imagine an incredibly smart ChatGPT: unhobbling gains should mean that this looks more like a drop-in remote worker, an incredibly smart agent that can reason and plan and error-correct and knows everything about you and your company and can work on a problem independently for weeks.
我们正朝着 2027 年实现 AGI(通用人工智能)的目标前进。这些 AI 系统基本上将能够自动化所有认知工作(想想:所有可以远程完成的工作)。
We are on course for AGI by 2027. These AI systems will basically be able to automate basically all cognitive jobs (think: all jobs that could be done remotely).
需要明确的是——误差范围很大。如果我们用尽数据,且突破数据墙所需的算法突破比预期更难,进展可能会停滞。也许解禁不会走得太远,我们只能停留在专家级聊天机器人,而非专家级同事。也许持续十年的趋势线会断裂,或者深度学习扩展这次真的撞墙了。(或者,算法突破,甚至仅仅是释放测试时算力过剩的简单解禁,可能带来范式转变,进一步加速进展,导致 AGI 更早到来。)
To be clear—the error bars are large. Progress could stall as we run out of data, if the algorithmic breakthroughs necessary to crash through the data wall prove harder than expected. Maybe unhobbling doesn’t go as far, and we are stuck with merely expert chatbots, rather than expert coworkers. Perhaps the decade-long trendlines break, or scaling deep learning hits a wall for real this time. (Or an algorithmic breakthrough, even simple unhobbling that unleashes the test-time compute overhang, could be a paradigm-shift, accelerating things further and leading to AGI even earlier.)
无论如何,我们正在快速跨越数量级,要极其认真地对待 2027 年实现 AGI——真正的 AGI——的可能性,这不需要异端信仰,只需对直线进行趋势外推。
In any case, we are racing through the OOMs, and it requires no esoteric beliefs, merely trend extrapolation of straight lines, to take the possibility of AGI—true AGI—by 2027 _extremely_ seriously.
如今,似乎很多人都在玩向下定义 AGI 的游戏,比如仅仅将其视为一个非常好的聊天机器人或其他什么。我指的是一个能完全自动化我或我朋友工作的 AI 系统,能完全完成 AI 研究员或工程师的工作。也许某些领域,比如机器人技术,默认情况下可能需要更长时间才能解决。而社会层面的推广,例如在医疗或法律行业,很容易因社会选择或监管而放缓。但一旦模型能自动化 AI 研究本身,那就足够了——足以启动强烈的反馈循环——我们就能非常迅速地取得进一步进展,自动化 AI 工程师自己解决所有剩余瓶颈,最终完全自动化一切。特别是,数百万个自动化研究者很可能将十年的算法进步压缩到一年或更短。AGI 仅仅是将很快到来的超级智能的一个小小前奏。(更多内容将在下一篇文章中讨论。)
It seems like many are in the game of downward-defining AGI these days, as just as really good chatbot or whatever. What I mean is an AI system that could fully automate my or my friends’ job, that could fully do the work of an AI researcher or engineer. Perhaps some areas, like robotics, might take longer to figure out by default. And the societal rollout, e.g. in medical or legal professions, could easily be slowed by societal choices or regulation. But once models can automate AI research itself, that’s enough—enough to kick off intense feedback loops—and we could very quickly make further progress, the automated AI engineers themselves solving all the remaining bottlenecks to fully automating everything. In particular, millions of automated researchers could very plausibly compress a decade of further algorithmic progress into a year or less. AGI will merely be a small taste of the superintelligence soon to follow. (More on that in the next piece.)
无论如何,不要指望这种令人眩晕的进展速度会放缓。趋势线看似无害,但其影响却极为强烈。如同之前的每一代模型,每一代新模型都会让大多数旁观者目瞪口呆;当模型很快解决需要博士花费数天的极其困难的科学问题时,当它们在你的电脑上飞速完成你的工作时,当它们从头编写数百万行代码的代码库时,当这些模型每一年或两年创造的经济价值增长 10 倍时,他们会难以置信。忘记科幻,数一数数量级:这正是我们应期待的。AGI 不再是遥远的幻想。扩展简单的深度学习技术一直有效,模型只想学习,而到 2027 年底,我们将再次实现超过 10 万倍的提升。用不了多久,它们就会比我们更聪明。
In any case, do not expect the vertiginous pace of progress to abate. The trendlines look innocent, but their implications are intense. As with every generation before them, every new generation of models will dumbfound most onlookers; they’ll be incredulous when, very soon, models solve incredibly difficult science problems that would take PhDs days, when they’re whizzing around your computer doing your job, when they’re writing codebases with millions of lines of code from scratch, when every year or two the economic value generated by these models 10xs. Forget scifi, count the OOMs: it’s what we should expect. AGI is no longer a distant fantasy. Scaling up simple deep learning techniques has just worked, the models just want to learn, and we’re about to do another 100,000x+ by the end of 2027. It won’t be long before they’re smarter than us.
_GPT-4 只是开始——四年后我们会在哪里?不要犯低估深度学习快速进展的错误(如 GANs 进展所示)。_
_GPT-4 is just the beginning—where will we be four years later? Do not make the mistake of underestimating the rapid pace of deep learning progress (as illustrated byprogress in GANs)._
_II. 从 AGI 到超级智能:智能爆炸_
_II. From AGI to Superintelligence: the Intelligence Explosion_
我曾对 AGI 的短时间线持怀疑态度。一个原因是,将如此多的 AGI 概率质量集中在这十年似乎不合理(这看起来像是典型的“我们如此特别”的谬误)。我认为我们应该对实现 AGI 所需的条件保持不确定,这应该会导致一个更加“分散”的 AGI 出现时间概率分布。
I used to be more skeptical of short timelines to AGI. One reason is that it seemed unreasonable to privilege this decade, concentrating so much AGI-probability-mass on it (it seemed like a classic fallacy to think “oh we’re so special”). I thought we should be uncertain about what it takes to get AGI, which should lead to a much more “smeared-out” probability distribution over when we might get AGI.
然而,我改变了想法:关键在于,我们对实现 AGI 所需条件的不确定性应该体现在“数量级”(有效算力)上,而不是年份上。
However, I’ve changed my mind: critically, our uncertainty over what it takes to get AGI should be over _OOMs_ (of effective compute), rather than over years.
我们正在这十年间快速跨越数量级。即使在摩尔定律的黄金时代,它也只是每十年 1-1.5 个数量级。我估计我们将在 4 年内完成约 5 个数量级,整个十年内完成超过 10 个数量级。
We’re racing through the OOMs this decade. Even at its bygone heyday, Moore’s law was only 1–1.5 OOMs/decade. I estimate that we will do ~5 OOMs in 4 years, and over ~10 this decade overall.
我们在这十年间一直在快速跨越数量级;在 2030 年代初之后,我们将面临缓慢的跋涉。
_We’ve been racing through the OOMs this decade; after the early 2030s, we will face a slow slog._
本质上,我们正处于一次大规模扩张的中期,这十年间正在收获一次性的收益,此后通过数量级的进展将慢得多。如果这次扩张未能在未来 5-10 年内让我们达到 AGI,那么可能还需要很长时间。
In essence, we’re in the middle of a huge scaleup reaping one-time gains this decade, and progress through the OOMs will be multiples slower thereafter. If this scaleup doesn’t get us to AGI in the next 5-10 years, it might be a long way out.
* _支出扩张_:过去花一百万美元训练一个模型是离谱的;到本十年末,我们很可能拥有 1000 亿美元或 1 万亿美元的集群。再往上走将很困难;这基本上已经是可行极限(无论是从大企业能负担得起的角度,还是仅作为 GDP 的一部分)。此后,我们只能依靠每年 2%的缓慢实际 GDP 增长趋势来增加这一数字。
* _Spending scaleup_: Spending a million dollars on a model used to be outrageous; by the end of the decade, we will likely have $100B or $1T clusters. Going much higher than that will be hard; that’s already basically the feasible limit (both in terms of what big business can afford, and even just as a fraction of GDP). Thereafter all we have is glacial 2%/year trend real GDP growth to increase this.
* _硬件收益_:AI 硬件的改进速度远快于摩尔定律。这是因为我们一直在为 AI 工作负载定制芯片。例如,我们从 CPU 转向 GPU;为 Transformer 适配芯片;并且我们采用了更低精度的数字格式,从传统超级计算的 fp64/fp32 到 H100 上的 fp8。这些都是巨大的收益,但到本十年末,我们很可能拥有完全专门化的 AI 专用芯片,而进一步超越摩尔定律的收益可能有限。
* _Hardware gains_: AI hardware has been improving much more quickly than Moore’s law. That’s because we’ve been specializing chips for AI workloads. For example, we’ve gone from CPUs to GPUs; adapted chips for Transformers; and we’ve gone down to much lower precision number formats, from fp64/fp32 for traditional supercomputing to fp8 on H100s. These are large gains, but by the end of the decade we’ll likely have totally-specialized AI-specific chips, without much further beyond-Moore’s law gains possible.
* _算法进步_:在未来十年,AI 实验室将在算法研发上投入数百亿美元,世界上最聪明的人都将致力于此;从微小的效率提升到新的范式,我们将摘取许多低垂的果实。我们可能不会达到任何硬性限制(尽管“解绑”可能是有限的),但至少改进的速度应该会放缓,因为快速增长(在资金和人力资本投资方面)必然会放缓(例如,大多数聪明的 STEM 人才已经将从事 AI 工作)。(也就是说,这是最不确定的预测,也是上图中 2030 年代数量级不确定性的主要来源。)
* _Algorithmic progress_: In the coming decade, AI labs will invest tens of billions in algorithmic R&D, and all the smartest people in the world will be working on this; from tiny efficiencies to new paradigms, we’ll be picking lots of the low-hanging fruit. We probably won’t reach any sort of hard limit (though “unhobblings” are likely finite), but at the very least the pace of improvements should slow down, as the rapid growth (in $ and human capital investments) necessarily slows down (e.g., most of the smart STEM talent will already be working on AI). (That said, this is the most uncertain to predict, and the source of most of the uncertainty on the OOMs in the 2030s on the plot above.)
综合来看,这意味着我们在下一个十年跨越的数量级比之后几十年可能跨越的还要多。也许这足够了——我们很快就能实现 AGI——或者我们可能面临漫长而缓慢的跋涉。你我完全可以对 AGI 的_中位数_时间持有不同意见,这取决于我们认为实现 AGI 有多难——但鉴于我们目前正在快速跨越数量级,你的_众数_AGI 年份肯定应该在十年后的某个时候。
Put together, this means we are racing through many more OOMs in the next decade than we might in multiple decades thereafter. Maybe it’s enough—and we get AGI soon—or we might be in for a long, slow slog. You and I can reasonably disagree on the _median_ time to AGI, depending on how hard we think achieving AGI will be—but given how we’re racing through the OOMs right now, certainly your _modal_ AGI year should sometime later this decade or so.
Matthew Barnett 有一个很好的相关可视化,只考虑了计算和生物限制。
Matthew Barnett has a nice related visualization of this, considering just compute and biological bounds.
1. 他们过去十年每年都做出预测,并且一直错误……↩
1. Predictions they’ve made every year for the last decade, and which they’ve been consistently wrong about…↩
2. 来自 SSC:Janelle Shane 问 GPT-2 它最喜欢的十种动物:
2. From SSC: Janelle Shane asks GPT-2 its ten favorite animals:
翼展约 4 英寸的刀嘴海雀,以及青蛙身上的心形纹身
Razorbill with wings hanging about 4 inches from one’s face and a heart tattoo on a frog
可以致盲、切割并生吃的鸡蛇互锁四足动物:
Cockatric interlocking tetrabods that can be blind, cut, and eaten raw:
生活在阳光下的黑白沙漠鳄鱼
Black and white desert crocodiles living in sunlight
4. 我指的是笨重的旧版 GPT-3,而不是你可能从 ChatGPT 知道的显著改进的 GPT-3.5。↩
4. I mean clunky old GPT-3 here, not the dramatically-improved GPT-3.5 you might know from ChatGPT.↩
5. 不,这些测试不在训练集中。AI 实验室确实努力确保这些评估不受污染,因为他们需要良好的测量来进行良好的科学。ScaleAI 最近的一项分析证实,领先的实验室并没有过度拟合基准(尽管一些较小的 LLM 开发者可能正在夸大他们的数字)。↩
5. And no, these tests aren’t in the training set. AI labs put real effort into ensuring these evals are uncontaminated, because they need good measurements in order to do good science. A recent analysis on this by ScaleAI confirmed that the leading labs aren’t overfitting to the benchmarks (though some smaller LLM developers might be juicing their numbers).↩
6. 在原始论文中,指出:“我们还评估了人类在 MATH 上的表现,发现一个不太喜欢数学的计算机科学博士生在 MATH 上获得了约 40%的分数,而一位三次 IMO 金牌得主获得了 90%,这表明 MATH 对人类来说也可能具有挑战性。”↩
6. In the original paper, it was noted: “We also evaluated humans on MATH, and found that a computer science PhD student who does not especially like mathematics attained approximately 40% on MATH, while a three-time IMO gold medalist attained 90%, indicating that MATH can be challenging for humans as well.”↩
7. 一位合著者指出:“当我们的团队首次发布 MATH 数据集时,至少有一位[机器学习研究员同事]告诉我们,这是一个无意义的数据集,因为它远远超出了 ML 模型所能完成的范围(事实上,我自己也有点担心这一点)。”↩
7. A coauthor notes: “When our group first released the MATH dataset, at least one [ML researcher colleague] told us that it was a pointless dataset because it was too far outside the range of what ML models could accomplish (indeed, I was somewhat worried about this myself).”↩
8. 这是 Yann LeCun 在 2022 年预测,即使 GPT-5000 也无法推理与现实世界的物理交互;一年后 GPT-4 显然轻松做到了。
8. Here’s Yann LeCun predicting in 2022 that even GPT-5000 won’t be able to reason about physical interactions with the real world; GPT-4 obviously does it with ease a year later.
这是 Gary Marcus 在 GPT-2 之后预测的墙被 GPT-3 解决,以及他在 GPT-3 之后预测的墙被 GPT-4 解决。
Here’s Gary Marcus’s walls predicted after GPT-2 being solved by GPT-3, and the walls he predicted after GPT-3 being solved by GPT-4.
这是 Bryan Caplan 教授输掉了他有史以来的第一次公开赌注(此前他以完美的公开赌注记录而闻名)。2023 年 1 月,在 GPT-3.5 在他的经济学期中考试中获得 D 后,Caplan 教授与 Matthew Barnett 打赌,到_2029 年_之前没有 AI 能在他的经济学期中考试中获得 A。仅仅两个月后,当 GPT-4 发布时,它立即在他的期中考试中获得了 A(并且这将是班上最高的分数之一)。↩
Here’s Prof. Bryan Caplan losing his first-ever public bet (after previously famously having a perfect public betting track record). In January 2023, after GPT-3.5 got a D on his economics midterm, Prof. Caplan bet Matthew Barnett that no AI would get an A on his economics midterms by _2029._ Just two months later, when GPT-4 came out, it promptly scored an A on his midterm (and it would have been one of the highest scores in his class).↩
9. 在钻石集上,模型使用思维链尝试 32 次后的多数投票。↩
9. On the diamond set, majority voting of the model trying 32 times with chain-of-thought.↩
10. 值得注意的是这些趋势线的一致性。将原始缩放定律论文与自那以来的算力和算力效率缩放的一些估计相结合,意味着在超过 15 个数量级(超过 1,000,000,000,000,000 倍有效算力)上存在一致的缩放趋势。
10. And it’s worth noting just how consistent these trendlines are. Combining the original scaling laws paper with some of the estimates on compute and compute efficiency scaling since then implies a consistent scaling trend for over 15 orders of magnitude (over 1,000,000,000,000,000x in effective compute)
11. 一个常见的误解是缩放仅适用于困惑度损失,但我们在基准测试的下游性能上也看到了非常清晰和一致的缩放行为。通常只需找到正确的对数-对数图。例如,在 GPT-4 博客文章中,他们展示了使用 MLPR(平均对数通过率)在 6 个数量级(1,000,000 倍)的算力上,编码问题性能的一致缩放行为。“涌现能力是海市蜃楼吗?”论文也提出了类似的观点;通过选择正确的度量,下游任务的性能几乎总是存在一致的趋势。
11. A common misconception is that scaling only holds for perplexity loss, but we see very clear and consistent scaling behavior on downstream performance on benchmarks as well. It’s usually just a matter of finding the right log-log graph. For example, in the GPT-4 blog post, they show consistent scaling behavior for performance on coding problems over 6 OOMs (1,000,000x) of compute, using MLPR (mean log pass rate). The “Are Emergent Abilities a Mirage?” paper makes a similar point; with the right choice of metric, there is almost always a consistent trend for performance on downstream tasks.
更一般地说,“缩放假设”的定性观察——模型能力随规模变化的非常清晰的趋势——早于损失缩放曲线;“缩放定律”工作只是对此的正式测量。
More generally, the “scaling hypothesis” qualitative observation—very clear trends on model capability with scale—predates loss-scaling-curves; the “scaling laws” work was just a formal measurement of this.
12. 1. Gemini 1.5 Flash 在 MATH 上得分为 54.9%,成本为每百万 token $0.35/$1.05(输入/输出)。GPT-4 在发布前 MATH 得分为 42.5%,2023 年初 MATH 得分为 52.9%,成本为每百万 token $30/$60(输入/输出);这比 Gemini 1.5 Flash 每个 token 贵 85 倍/57 倍(输入/输出)。为了保守起见,我使用上述 30 倍成本降低的估计(考虑到 Gemini 1.5 Flash 可能使用更多 token 来推理问题)。
12. 1. Gemini 1.5 Flash scores 54.9% on MATH, and costs $0.35/$1.05 (input/output) per million tokens. GPT-4 scored 42.5% on MATH prelease and 52.9% on MATH in early 2023, and cost $30/$60 (input/output) per million tokens; that’s 85x/57x (input/output) more expensive per token than Gemini 1.5 Flash. To be conservative, I use an estimate of 30x cost decrease above (accounting for Gemini 1.5 Flash possibly using more tokens to reason through problems).
2. Minerva540B 在 MATH 上得分为 50.3%,使用 64 个样本的多数投票。一位知情朋友估计,这里的基础模型推理成本可能比 GPT-4 贵 2-3 倍。然而,Minerva 在快速抽查中似乎每个答案使用的 token 较少。更重要的是,Minerva 需要 64 个样本才能达到该性能,这暗示如果通过推理 API 天真地运行,成本将增加 64 倍。在实践中,运行评估时可以缓存提示 token;给定少量提示,提示 token 可能占成本的大部分,即使考虑到输出 token。假设输出 token 占单个样本成本的三分之一,那么通过缓存进行 maj@64 只会导致成本增加约 20 倍。为了保守起见,我在上述中使用粗略的 20 倍成本降低(即使通过 API 运行此操作的天真推理成本降低会更大)。↩
2. Minerva540B scores 50.3% on MATH, using majority voting among 64 samples. A knowledgeable friend estimates the base model here is probably 2-3x more expensive to inference than GPT-4. However, Minerva seems to use somewhat fewer tokens per answer on a quick spot check. More importantly, Minerva needed 64 samples to achieve that performance, naively implying a 64x multiple on cost if you e.g. naively ran this via an inference API. In practice, prompt tokens can be cached when running an eval; given a few-shot prompt, prompt tokens are likely a majority of the cost, even accounting for output tokens. Supposing output tokens are a third of the cost for getting a single sample, that would imply only a ~20x increase in cost from the maj@64 with caching. To be conservative, I use the rough number of a 20x cost decrease in the above (even if the naive decrease in inference cost from running this via an API would be larger).↩
13. 尽管这些是推理效率(而非训练效率),并且在某种程度上将反映推理特定的优化,但 a)它们表明大量算法进步是可能的并且正在发生,b)通常算法改进既是训练效率提升也是推理效率提升,例如通过减少所需参数数量。↩
13. Though these are inference efficiencies (rather than necessarily training efficiencies), and to some extent will reflect inference-specific optimizations, a) they suggest enormous amounts of algorithmic progress is possible and happening in general, and b) it’s often the case that an algorithmic improvements is both a training efficiency gain and an inference efficiency, for example by reducing the number of parameters necessary.↩
14. GPT-3:$60/1M tokens,GPT-4:$30/1M 输入 tokens 和$60/1M 输出 tokens。↩
14. GPT-3: $60/1M tokens, GPT-4: $30/1M input tokens and $60/1M output tokens.↩
15. Chinchilla 缩放定律表明,参数数量和数据的缩放应相等。也就是说,参数数量增长的有效训练算力数量级的一半。同时,参数数量直观上大致与推理成本成正比。在其他条件相同的情况下,恒定的推理成本因此意味着有效算力增长的一半数量级被算法胜利“抵消”了。
15. Chinchilla scaling laws say that one should scale parameter count and data equally. That is, parameter count grows “half the OOMs” of the OOMs that effective training compute grows. At the same time, parameter count is intuitively roughly proportional to inference costs. All else equal, constant inference costs thus implies that half of the OOMs of effective compute growth were “canceled out” by algorithmic win.
也就是说,明确地说,这是一个非常粗略的计算(仅用于粗略说明),在各种方面都是错误的。可能存在推理特定的优化(不会转化为训练效率);可能存在训练效率提升但不减少参数数量(因此不转化为推理效率);等等。↩
That said, to be clear, this is a very naive calculation (just meant for a rough illustration) that is wrong in various ways. There may be inference-specific optimizations (that don’t translate into training efficiency); there may be training efficiencies that don’t reduce parameter count (and thus don’t translate into inference efficiency); and so on.↩
16. Gemini 1.5 Flash 在 LMSys(一个聊天机器人排行榜)上的排名与 GPT-4 相似(高于原始 GPT-4,低于 GPT-4 的更新版本),并且在 MATH 和 GPQA(衡量推理能力的评估)上的表现与原始 GPT-4 相似,同时在 MMLU(更侧重于衡量知识的评估)上大致介于 GPT-3.5 和 GPT-4 之间。↩
16. Gemini 1.5 Flash ranks similarly to GPT-4 (higher than original GPT-4, lower than updated versions of GPT-4) on LMSys, a chatbot leaderboard, and has similar performance on MATH and GPQA (evals that measure reasoning) as the original GPT-4, while landing roughly in the middle between GPT-3.5 and GPT-4 on MMLU (an eval that more heavily weights towards measuring knowledge).↩
17. 在~GPT-3 规模下,超过 3 倍;在更大规模下更多。↩
17. At ~GPT-3 scale, more than 3x at larger scales.↩
18. 例如,本文包含了对 GPT-3 风格 vanilla Transformer 与多年来发布的各种简单架构和训练配方更改(RMSnorms 代替 layernorm,不同的位置嵌入,SwiGlu 激活,AdamW 优化器代替 Adam 等)的比较,他们称之为“Transformer++”,暗示至少在小规模下有 6 倍的增益。↩
18. For example, this paper contains a comparison of a GPT-3-style vanilla Transformer to various simple changes to architecture and training recipe published over the years (RMSnorms instead of layernorm, different positional embeddings, SwiGlu activation, AdamW optimizer instead of Adam, etc.), what they call “Transformer++”, implying a 6x gain at least at small scale.↩
19. 如果我们采用 0.5 数量级/年的趋势,以及 GPT-2 和 GPT-4 发布之间的 4 年,那将是 2 个数量级。然而,GPT-2 到 GPT-3 是一个简单的规模扩张(在 Transformer 等带来的巨大收益之后),而 OpenAI 声称 GPT-4 预训练在 2022 年完成,这可能意味着我们在这里应该计算接近 2 年的算法进步。1 个数量级的算法效率似乎是一个保守的下限。↩
19. If we take the trend of 0.5 OOMs/year, and 4 years between GPT-2 and GPT-4 release, that would be 2 OOMs. However, GPT-2 to GPT-3 was a simple scaleup (after big gains from e.g. Transformers), and OpenAI claims GPT-4 pretraining finished in 2022, which could mean we’re looking at closer to 2 years worth of algorithmic progress that we should be counting here. 1 OOM of algorithmic efficiency seems like a conservative lower bound.↩
20. 至少,考虑到十多年来一致的算法改进,举证责任将落在那些认为这一切会突然停止的人身上。
20. At the very least, given over a decade of consistent algorithmic improvements, the burden of proof would be on those who would suggest it will all suddenly come to a halt
21. 考虑到集群成本,3 倍算力效率的经济回报将以数百亿或更多美元衡量。↩
21. The economic returns to a 3x compute efficiency will be measured in the $10s of billions or more, given cluster costs.↩
22. 非常粗略地大约是~10 倍的增益。↩
22. Very roughly something like a ~10x gain.↩
23. 并且一遍又一遍地重读同一本教科书可能导致记忆,而不是理解。我认为这就是许多书呆子通过数学考试的方式。
23. And just rereading the same textbook over and over again might result in memorization, not understanding. I take it that’s how many wordcels pass math classes
24. 另一种我觉得有趣的思考方式:预训练和上下文学习之间存在“缺失的中间地带”。上下文学习是_不可思议的_(并且与人类样本效率竞争)。例如,Gemini 1.5 Pro 论文讨论了给模型提供关于 Kalamang 语(一种由不到 200 人使用且基本上不在互联网上存在的语言)的教学材料(教科书、词典)——模型学会了从英语翻译到 Kalamang 语,达到人类水平!在上下文中,模型能够像人类一样从教科书中学习(并且比仅仅将那一本教科书扔进预训练中学习得好得多)。
24. One other way of thinking about it I find interesting: there is a “missing-middle” between pretraining and in-context learning. In-context learning is _incredible_ (and competitive with human sample efficiency). For example, the Gemini 1.5 Pro paper discusses giving the model instructional materials (a textbook, a dictionary) on Kalamang, a language spoken by fewer than 200 people and basically not present on the internet, in context—and the model learns to translate from English to Kalamang at human-level! In context, the model is able to learn from the textbook as well as a human could (and much better than it would learn from just chucking that one textbook into pretraining).
当人类从教科书中学习时,他们能够通过练习将短期记忆/学习提炼为长期记忆/长期技能;然而,我们没有等效的方法将上下文学习“提炼回权重”。合成数据/自我对弈/强化学习等正在尝试解决这个问题:让模型自己学习,然后思考并练习所学内容,将该学习提炼回权重。↩
When a human learns from a textbook, they’re able to distill their short-term memory/learnings into long-term memory/long-term skills with practice; however, we don’t have an equivalent way to distill in-context learning “back to the weights.” Synthetic data/self-play/RL/etc are trying to fix that: let the model learn by itself, then think about it and practice what it learned, distilling that learning back into the weights.↩
25. 另见 Andrej Karpathy 在此讨论的演讲。↩
25. See also Andrej Karpathy’s talk discussing this here.↩
26. 在某种意义上,这就是无监督学习的魔力:为了更好地预测下一个词,降低困惑度,模型学习了极其丰富的内部表示,从(著名的)情感到复杂的世界模型。但是,开箱即用时,它们受到限制:它们仅仅使用其令人难以置信的内部表示来预测随机互联网文本中的下一个词,而不是以最佳方式应用它们来实际解决你的问题。↩
26. That’s the magic of unsupervised learning, in some sense: to better predict the next token, to make perplexity go down, models learn incredibly rich internal representations, everything from (famously) sentiment to complex world models. But, out of the box, they’re hobbled: they’re using their incredible internal representations merely to predict the next token in random internet text, and rather than applying them in the best way to actually try to solve your problem.↩
27. 参见更新的 Gemini 1.5 白皮书中的图 7,比较了 Gemini 1.5 Pro 和 Gemini 1.5 Flash(一个更便宜且可能更小的模型)的困惑度与上下文。↩
27. See Figure 7 from the updated Gemini 1.5 whitepaper, comparing perplexity vs. context for Gemini 1.5 Pro and Gemini 1.5 Flash (a much cheaper and presumably smaller model).↩
29. 这很合理——为什么它会学到更长视野的推理和纠错技能?互联网上很少有数据以“我完整的内心独白、推理、一个月内我从事项目时所有相关步骤”的形式存在。释放这种能力将需要一种新的训练,让它学习这些额外的技能。
29. Which makes sense—why would it have learned the skills for longer-horizon reasoning and error correction? There’s very little data on the internet in the form of “my complete internal monologue, reasoning, all the relevant steps over the course of a month as I work on a project.” Unlocking this capability will require a new kind of training, for it to learn these extra skills.
或者正如 Gwern 所说(私人通信):“‘星系大小的脑袋,他们让我做什么?预测基准上拼写错误的答案!’抑郁的神经网络 Marvin 呻吟道。”↩
Or as Gwern put it (private correspondence): “‘Brain the size of a galaxy, and what do they ask me to do? Predict the misspelled answers on benchmarks!’ Marvin the depressed neural network moaned.”↩
30. 系统 I 与系统 II 是思考当前 LLM 能力(包括其局限性和愚蠢错误)以及通过强化学习和解绑可能实现的目标的有用方式。这样想:当你开车时,大多数时候你处于自动驾驶状态(系统 I,模型现在主要做的事情)。但是当你遇到复杂的施工区或陌生的交叉路口时,你可能会让副驾驶暂停对话一会儿,同时你弄清楚——实际思考——发生了什么以及该怎么做。如果你被迫只依靠系统 I 生活(更接近今天的模型),你会遇到很多麻烦。创建系统 II 推理循环的能力是一个核心解锁。↩
30. System I vs. System II is a useful way of thinking about current capabilities of LLMs—including their limitations and dumb mistakes—and what might be possible with RL and unhobbling. Think of this way: when you are driving, most of the time you are on autopilot (system I, what models mostly do right now). But when you encounter a complex construction zone or novel intersection, you might ask your passenger-seat-companion to pause your conversation for a moment while you figure out—actually think about—what’s going on and what to do. If you were forced to go about life with only system I (closer to models today), you’d have a lot of trouble. Creating the ability for system II reasoning loops is a central unlock.↩
31. 基于上述关于物理算力和算法效率规模扩张的最佳猜测假设,并简化并行性考虑(实际上,它可能看起来更像“一天 1440(60*24)个 GPT-4 级别模型”或类似)。↩
31. On the best guess assumptions on physical compute and algorithmic efficiency scaleups described above, and simplifying parallelism considerations (in reality, it might look more like “1440 (60*24) GPT-4-level models in a day” or similar).↩
32. 当然,我们今天拥有的任何基准都会被饱和。但这说明不了什么;它主要反映了制作足够困难基准的难度。↩
32. Of course, any benchmark we have today will be saturated. But that’s not saying much; it’s mostly a reflection on the difficulty of making hard-enough benchmarks.↩