Why I don't think AGI is right around the corner
打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→通过订阅,您同意 Substack 的使用条款,并确认其信息收集通知和隐私政策。
By subscribing, you agree Substack's Terms of Use, and acknowledge its Information Collection Notice and Privacy Policy.
订阅即表示您同意 Substack 的使用条款,并确认其信息收集声明和隐私政策。
By subscribing, you agree Substack's Terms of Use, and acknowledge its Information Collection Notice and Privacy Policy.
“事情发生的时间比你想象的要长,然后它们发生的速度比你想象的要快。”——鲁迪格·多恩布什
“Things take longer to happen than you think they will, and then they happen faster than you thought they could.” - Rudiger Dornbusch
我在播客中进行了很多讨论,争论 AGI 的时间线。一些嘉宾认为还有 20 年,另一些则认为只有 2 年。以下是我截至 2025 年 6 月的想法。
I’ve had a lot of discussions on my podcast where we haggle out timelines to AGI. Some guests think it’s 20 years away - others 2 years. Here’s where my thoughts stand as of June 2025.
有时人们说,即使所有 AI 进展完全停止,今天的系统在经济上的变革性仍将远超互联网。我不同意。我认为今天的 LLM 很神奇。但财富 500 强公司没有用它们来改造工作流程,原因并非管理层过于保守。相反,我认为让 LLM 提供正常的人类式劳动确实很难。这与这些模型缺乏的一些基本能力有关。
Sometimes people say that even if all AI progress totally stopped, the systems of today would still be far more economically transformative than the internet. I disagree. I think the LLMs of today are magical. But the reason that the Fortune 500 aren’t using them to transform their workflows isn’t because the management is too stodgy. Rather, I think it’s genuinely hard to get normal humanlike labor out of LLMs. And this has to do with some fundamental capabilities these models lack.
我喜欢认为自己在 Dwarkesh 播客中算是“AI 前沿”。我可能花了超过一百个小时尝试为我的后期制作流程构建小型 LLM 工具。试图让它们变得有用的经历延长了我的时间线。我尝试让 LLM 像人类一样重写自动生成的转录稿以提高可读性。或者我尝试让它们从转录稿中识别出片段用于推文。有时我尝试让它们与我逐段合写一篇文章。这些都是简单、独立、短周期、语言输入语言输出的任务——本应是 LLM 技能组合的核心。它们在这些任务上只能打 5/10 分。别误会,这已经很令人印象深刻了。
I like to think I’m “AI forward” here at the Dwarkesh Podcast. I’ve probably spent over a hundred hours trying to build little LLM tools for my post production setup. And the experience of trying to get them to be useful has extended my timelines. I’ll try to get the LLMs to rewrite autogenerated transcripts for readability the way a human would. Or I’ll try to get them to identify clips from the transcript to tweet out. Sometimes I’ll try to get them to co-write an essay with me, passage by passage. These are simple, self contained, short horizon, language in-language out tasks - the kinds of assignments that should be dead center in the LLMs’ repertoire. And they're 5/10 at them. Don’t get me wrong, that’s impressive.
但根本问题是,LLM 不会像人类那样随着时间的推移变得更好。缺乏持续学习是一个巨大的问题。LLM 在许多任务上的基线可能高于普通人。但无法给模型提供高层次反馈。你只能使用开箱即用的能力。你可以不断调整系统提示。实际上,这根本无法产生与人类员工所经历的学习和改进相媲美的效果。
But the fundamental problem is that LLMs don’t get better over time the way a human would. The lack of continual learning is a huge huge problem. The LLM baseline at many tasks might be higher than an average human's. But there’s no way to give a model high level feedback. You’re stuck with the abilities you get out of the box. You can keep messing around with the system prompt. In practice this just doesn’t produce anything even close to the kind of learning and improvement that human employees experience.
人类之所以如此有用,主要不是因为他们原始的智力。而是因为他们能够积累上下文、审视自己的失败,并在练习任务时获得小的改进和效率提升。
The reason humans are so useful is not mainly their raw intelligence. It’s their ability to build up context, interrogate their own failures, and pick up small improvements and efficiencies as they practice a task.
如何教一个孩子吹萨克斯?你让她试着吹一下,听声音如何,然后调整。现在想象用这种方式教萨克斯:一个学生尝试一次。一旦她犯错,你就把她打发走,并写下关于哪里出错的详细说明。下一个学生阅读你的笔记,然后尝试直接演奏 Charlie Parker。当她失败时,你为下一个学生完善说明。
How do you teach a kid to play a saxophone? You have her try to blow into one, listen to how it sounds, and adjust. Now imagine teaching saxophone this way instead: A student takes one attempt. The moment they make a mistake, you send them away and write detailed instructions about what went wrong. The next student reads your notes and tries to play Charlie Parker cold. When they fail, you refine the instructions for the next student.
这根本行不通。无论你的提示打磨得多么好,没有一个孩子能仅仅通过阅读你的说明就学会吹萨克斯。但这是我们作为用户“教”LLM 的唯一方式。
This just wouldn’t work. No matter how well honed your prompt is, no kid is just going to learn how to play saxophone from just reading your instructions. But this is the only modality we as users have to ‘teach’ LLMs anything.
是的,有 RL 微调。但它不像人类学习那样是一个有意识的、自适应的过程。我的编辑变得非常出色。如果我们必须为他们工作中的不同子任务构建定制的 RL 环境,他们就不会变得这么好。他们自己注意到了很多小细节,并深入思考了什么能引起观众共鸣、什么内容能让我兴奋,以及如何改进他们的日常工作流程。
Yes, there’s RL fine tuning. But it’s just not a deliberate, adaptive process the way human learning is. My editors have gotten extremely good. And they wouldn’t have gotten that way if we had to build bespoke RL environments for different subtasks involved in their work. They’ve just noticed a lot of small things themselves and thought hard about what resonates with the audience, what kind of content excites me, and how they can improve their day to day workflows.
现在,可以想象某种方式,一个更智能的模型可以为自己构建一个专用的 RL 循环,从外部看起来非常自然。我给出一些高层次反馈,模型就会提出一堆可验证的练习题来进行 RL——甚至可能是一个完整的环境来练习它认为缺乏的技能。但这听起来真的很难。我不知道这些技术如何很好地泛化到不同类型的任务和反馈。最终,模型将能够以人类那种微妙自然的方式在工作中学习。然而,我很难看到这如何在未来几年内发生,因为目前没有明显的方法将在线、持续学习融入 LLM 这类模型中。
Now, it’s possible to imagine some way in which a smarter model could build a dedicated RL loop for itself which just feels super organic from the outside. I give some high level feedback, and the model comes up with a bunch of verifiable practice problems to RL on - maybe even a whole environment in which to rehearse the skills it thinks it's lacking. But this just sounds really hard. And I don’t know how well these techniques will generalize to different kinds of tasks and feedback. Eventually the models will be able to learn on the job in the subtle organic way that humans can. However, it’s just hard for me to see how that could happen within the next few years, given that there’s no obvious way to slot in online, continuous learning into the kinds of models these LLMs are.
LLM 实际上在会话过程中会变得有点聪明和有用。例如,有时我会与 LLM 合写一篇文章。我给它一个大纲,让它逐段起草文章。直到第 4 段之前,它的所有建议都很糟糕。所以我从头重写整个段落,并告诉它:“嘿,你的东西很烂。这是我写的。”这时,它实际上可以开始为下一段提供好的建议。但这种对我偏好和风格的微妙理解在会话结束时就会丢失。
LLMs actually do get kinda smart and useful in the middle of a session. For example, sometimes I’ll co-write an essay with an LLM. I’ll give it an outline, and I’ll ask it to draft the essay passage by passage. All its suggestions up till 4 paragraphs in will be bad. So I'll just rewrite the whole paragraph from scratch and tell it, "Hey, your shit sucked. This is what I wrote instead." At that point, it can actually start giving good suggestions for the next paragraph. But this whole subtle understanding of my preferences and style is lost by the end of the session.
也许解决这个问题的简单方法是一个长的滚动上下文窗口,就像 Claude Code 那样,每 30 分钟将会话内存压缩成一个摘要。我只是觉得,将所有这些丰富的隐性经验滴定成文本摘要,在软件工程(非常基于文本)之外的领域会很脆弱。再次想想用一长串关于你所学到的文本摘要来教某人吹萨克斯的例子。即使是 Claude Code,也常常会撤销我们之前一起努力获得的优化——因为做出优化的原因没有进入摘要。
Maybe the easy solution to this looks like a long rolling context window, like Claude Code has, which compacts the session memory into a summary every 30 minutes. I just think that titrating all this rich tacit experience into a text summary will be brittle in domains outside of software engineering (which is very text-based). Again, think about the example of trying to teach someone how to play the saxophone using a long text summary of your learnings. Even Claude Code will often reverse a hard-earned optimization that we engineered together before I hit /compact - because the explanation for why it was made didn’t make it into the summary.
这就是为什么我不同意 Sholto 和 Trenton 在我的播客上说的某些话(这句话来自 Trenton):
This is why I disagree with something Sholto and Trenton said on my podcast (this quote is from Trenton):
如果今天 AI 进展完全停滞,我认为<25%的白领工作会消失。当然,许多_任务_会被自动化。Claude 4 Opus 技术上可以为我重写自动生成的转录稿。但由于我无法让它随时间改进并学习我的偏好,我仍然雇人来做这件事。即使我们获得更多数据,如果没有持续学习的进展,我认为我们在白领工作方面将处于基本相似的境地——是的,技术上 AI 可能能够相当令人满意地完成许多子任务,但它们无法积累上下文,这将使它们无法作为实际员工在你的公司工作。
If AI progress totally stalls today, I think <25% of white collar employment goes away. Sure, many _tasks_ will get automated. Claude 4 Opus can technically rewrite autogenerated transcripts for me. But since it’s not possible for me to have it improve over time and learn my preferences, I still hire a human for this. Even if we get more data, without progress in continual learning, I think we will be in a substantially similar position with white collar work - yes, technically AIs might be able to do a lot of subtasks somewhat satisfactorily, but their inability to build up context will make it impossible to have them operate as actual employees at your firm.
虽然这让我对未来几年的变革性 AI 持悲观态度,但让我对未来几十年的 AI 特别乐观。当我们解决持续学习时,我们将看到模型价值的巨大不连续性。即使没有纯软件奇点(模型快速构建越来越智能的继任系统),我们可能仍然会看到类似广泛部署的智能爆炸。AI 将通过经济广泛部署,做不同的工作,并在工作中像人类一样学习。但与人类不同,这些模型可以整合所有副本的学习成果。因此,一个 AI 基本上在学习如何做世界上每一项工作。一个能够在线学习的 AI 可能在没有进一步算法进步的情况下迅速成为超级智能。
While this makes me bearish on transformative AI in the next few years, it makes me especially bullish on AI over the next decades. When we do solve continuous learning, we’ll see a huge discontinuity in the value of the models. Even if there isn’t a software only singularity (with models rapidly building smarter and smarter successor systems), we might still see something that looks like a broadly deployed intelligence explosion. AIs will be getting broadly deployed through the economy, doing different jobs and learning while doing them in the way humans can. But unlike humans, these models can amalgamate their learnings across all their copies. So one AI is basically learning how to do every single job in the world. An AI that is capable of online learning might functionally become a superintelligence quite rapidly without any further algorithmic progrss
然而,我不期望看到某个 OpenAI 直播宣布持续学习已完全解决。因为实验室有动力快速发布任何创新,我们会在看到真正像人类一样学习的东西之前,先看到一个有缺陷的早期版本的持续学习(或测试时训练——随便你怎么称呼)。我预计在这个大瓶颈完全解决之前,会得到很多预警。
However, I’m not expecting to watch some OpenAI livestream where they announce that continual learning has been totally solved. Because labs are incentivized to release any innovations quickly, we’ll see a broken early version of continual learning (or test time training - whatever you want to call it) before we see something which truly learns like a human. I expect to get lots of heads up before this big bottleneck is totally solved.
当我在播客中采访 Anthropic 研究员 Sholto Douglas 和 Trenton Bricken 时,他们表示预计明年年底前会出现可靠的计算机使用智能体。我们现在已经有计算机使用智能体,但它们还很差劲。他们设想的场景截然不同。他们的预测是,到明年年底,你应该能够告诉 AI“去帮我报税”。然后它会浏览你的电子邮件、亚马逊订单和 Slack 消息,与所有你需要发票的人来回发邮件,整理所有收据,判断哪些是业务支出,在边缘案例上征求你的批准,最后向 IRS 提交 1040 表格。
When I interviewed Anthropic researchers Sholto Douglas and Trenton Bricken on my podcast, they said that they expect reliable computer use agents by the end of next year. We already have computer use agents right now, but they’re pretty bad. They’re imagining something quite different. Their forecast is that by the end of next year, you should be able to tell an AI, “Go do my taxes.” And it goes through your email, Amazon orders, and Slack messages, emails back and forth with everyone you need invoices from, compiles all your receipts, decides which are business expenses, asks for your approval on the edge cases, and then submits Form 1040 to the IRS.
我对此持怀疑态度。我不是 AI 研究员,所以在技术细节上不敢反驳他们。但根据我所知的一点信息,以下是我押注这一预测不会成真的理由:
I’m skeptical. I’m not an AI researcher, so far be it for me to contradict them on technical details. But given what little I know, here’s why I’d bet against this forecast:
* 随着任务时间跨度的增加,展开过程必须变得更长。AI 需要完成两小时的智能体式计算机使用任务,我们才能判断它是否做对了。更不用说计算机使用需要处理图像和视频,这本身就已经更加消耗算力,即使不考虑更长的展开过程。这似乎会拖慢进展。
* As horizon lengths increase, rollouts have to become longer. The AI needs to do two hours worth of agentic computer use tasks before we can even see if it did it right. Not to mention that computer use requires processing images and video, which is already more compute intensive, even if you don’t factor in the longer rollout. This seems like this should slow down progress.
* 我们没有大规模的多模态计算机使用数据的预训练语料库。我喜欢 Mechanize 关于自动化软件工程的文章中的这句话:“在过去的十年 Scaling 中,我们被大量免费可用的互联网数据宠坏了。这些数据足以攻克自然语言处理,但不足以让模型成为可靠、有能力的智能体。想象一下,试图用 1980 年可用的所有文本数据来训练 GPT-4——即使我们有必要的算力,数据也远远不够。”
* We don’t have a large pretraining corpus of multimodal computer use data. I like this quote from Mechanize’s post on automating software engineering: “For the past decade of scaling, we’ve been spoiled by the enormous amount of internet data that was freely available for us to use. This was enough for cracking natural language processing, but not for getting models to become reliable, competent agents. Imagine trying to train GPT-4 on all the text data available in 1980—the data would be nowhere near enough, even if we had the necessary compute.”
再次强调,我不在实验室工作。也许纯文本训练已经为你提供了关于不同 UI 如何工作以及不同组件之间关系的良好先验。也许强化学习微调在样本效率上如此之高,以至于你不需要那么多数据。但我没有看到任何公开证据让我认为这些模型突然变得不那么数据饥渴,尤其是在这个它们经验明显不足的领域。
Again, I’m not at the labs. Maybe text only training already gives you a great prior on how different UIs work, and what the relationship between different components is. Maybe RL fine tuning is so sample efficient that you don’t need that much data. But I haven’t seen any public evidence which makes me think that these models have suddenly gotten less data hungry, especially in this domain where they’re substantially less practiced.
或者,也许这些模型是如此优秀的前端编码器,以至于它们可以为自己生成数百万个玩具 UI 来练习。对于我的反应,请参见下面的要点。
Alternatively, maybe these models are such good front end coders that they can just generate millions of toy UIs for themselves to practice on. For my reaction to this, see bullet point below.
* 即使是事后看来相当简单的算法创新,似乎也需要很长时间才能完善。DeepSeek 在其 R1 论文中解释的强化学习过程在高层次上看起来很简单。然而,从 GPT-4 发布到 o1 发布花了两年时间。当然,我知道说 R1/o1 很容易是极其傲慢的——需要大量的工程、调试和修剪替代想法才能得出这个解决方案。但这正是我的观点!看到实现“训练模型解决可验证的数学和编程问题”这个想法花了多长时间,让我觉得我们低估了解决更棘手的计算机使用问题的难度,在这个问题上,你是在完全不同的模态下操作,数据也少得多。
* Even algorithmic innovations which seem quite simple in retrospect seem to take a long time to iron out. The RL procedure which DeepSeek explained in their R1 paper seems simple at a high level. And yet it took 2 years from the launch of GPT-4 to the release of o1. Now of course I know it is hilariously arrogant to say that R1/o1 were easy - a ton of engineering, debugging, pruning of alternative ideas was required to arrive at this solution. But that’s precisely my point! Seeing how long it took to implement the idea, ‘Train the model to solve verifiable math and coding problems’, makes me think that we’re underestimating the difficulty of solving the much gnarlier problem of computer use, where you’re operating in a totally different modality with much less data.
好了,冷水泼够了。我不会像 Hackernews 上那些被宠坏的孩子一样,就算拿到一只会下金蛋的鹅,也只会抱怨它的叫声太吵。
Okay, enough cold water. I’m not going to be like one of those spoiled children on Hackernews who could be handed a golden-egg laying goose and still spend all their time complaining about how loud its quacks are.
你读过 o3 或 Gemini 2.5 的推理轨迹吗?那真的是在推理!它在分解问题,思考用户想要什么,对自己的内心独白做出反应,并在发现自己在追求无效方向时自我纠正。我们怎么能只是说:“哦,当然,机器会思考一堆,想出一堆点子,然后带回一个聪明的答案。机器就是干这个的。”
Have you read the reasoning traces of o3 or Gemini 2.5? It’s actually reasoning! It’s breaking down a problem, thinking through what the user wants, reacting to its own internal monologue, and correcting itself when it notices that it's pursuing an unproductive direction. How are we just like, “Oh yeah of course the machine is gonna go think a bunch, come up with a bunch of ideas, and come back with a smart answer. That’s what machines do.”
有些人过于悲观的部分原因是,他们没有在自己最擅长的领域里玩过最聪明的模型。给 Claude Code 一个模糊的规格说明,然后坐等 10 分钟,直到它零样本生成一个可用的应用程序,这是一种狂野的体验。它是怎么做到的?你可以谈论电路、训练分布、强化学习等等,但最直接、最简洁、最准确的解释就是:它是由婴儿般的通用智能驱动的。在这一点上,你心里一定有一部分在想:“它真的在起作用。我们正在制造有智能的机器。”
Part of the reason some people are too pessimistic is that they haven’t played around with the smartest models operating in the domains that they’re most competent in. Giving Claude Code a vague spec and sitting around for 10 minutes until it zero shots a working application is a wild experience. How did it do that? You could talk about circuits and the training distribution and RL and whatever, but the most proximal, concise, and accurate explanation is simply that it’s powered baby general intelligence. At this point, part of you has to be thinking, “It’s actually working. We’re making machines that are intelligent.”
我的概率分布非常宽。我想强调,我确实相信概率分布。这意味着为应对 2028 年可能出现的未对齐 ASI 做准备仍然很有意义——我认为这是一个完全可能的结果。
My probability distributions are super wide. And I want to emphasize that I do believe in probability distributions. Which means that work to prepare for misaligned 2028 ASI still makes a lot of sense - I think this is a totally plausible outcome.
但以下是我愿意以 50/50 概率打赌的时间线:
But here are the timelines where I’d take a 50/50 bet:
* AI 能像一位称职的总经理一样,在一周内端到端地为我处理小企业的税务:包括在各个网站上追查所有收据,找到所有缺失的部分,通过电子邮件与需要索取发票的人来回沟通,填写表格,并将其发送给 IRS:2028 年
* AI can do taxes end-to-end for my small business as well as a competent general manager could in a week: including chasing down all the receipts on different websites, finding all the missing pieces, emailing back and forth with anyone we need to hassle for invoices, filling out the form, and sending it to the IRS: 2028
* AI 在计算机使用方面能像人类一样轻松、自然、无缝且快速地学习,适用于任何白领工作。例如,如果我雇佣一个 AI 视频编辑,六个月后,它对我的偏好、我们的频道、什么对观众有效等有同样可操作的、深刻的理解,就像人类一样:2032 年
* I think we’re in the GPT 2 era for computer use. But we have no pretraining corpus, and the models are optimizing for a much sparser reward over a much longer time horizon using action primitives they’re unfamiliar with. That being said, the base model is decently smart and might have a good prior over computer use tasks, plus there’s a lot more compute and AI researchers in the world, so it might even out. Preparing taxes for a small business feels like for computer use what GPT 4 was for language. It took 4 years to get from GPT 2 to GPT 4.
你可能会反应说:“等等,你之前对持续学习这个障碍大做文章。但你的时间线却是,我们距离至少是广泛部署的智能爆炸只有 7 年。”是的,你说得对。我预测在相对短的时间内会出现一个相当疯狂的世界。
Just to clarify, I am not saying that we won’t have really cool computer use demos in 2026 and 2027 (GPT-3 was super cool, but not that practically useful). I’m saying that these models won’t be capable of end-to-end handling a week long and quite involved project which involves computer use.
AGI 时间线非常符合对数正态分布。要么是这十年,要么就没了。(不是真的没了,更像是每年的边际概率降低——但没那么吸引人。)过去十年 AI 的进步主要由前沿系统的训练算力 Scaling 驱动(每年超过 4 倍)。这不可能持续到这十年之后,无论你看芯片、电力,甚至用于训练的 GDP 比例。2030 年后,AI 进步必须主要来自算法进步。但即使在那里,低垂的果实也将被摘取(至少在深度学习范式下)。所以 AGI 的年度概率会骤降。
* AI learns on the job as easily, organically, seamlessly, and quickly as a human, for any white collar work. For example, if I hire an AI video editor, after six months, it has as much actionable, deep understanding of my preferences, our channel, what works for the audience, etc as a human would: 2032
这意味着,如果我们最终处于我 50/50 赌注中较长的那一边,我们很可能看到一个相对正常的世界,直到 2030 年代甚至 2040 年代。但在所有其他世界中,即使我们对 AI 当前的局限性保持清醒,也必须预期一些真正疯狂的结果。
* While I don’t see an obvious way to slot in continuous online learning into current models, 7 years is a long time! GPT 1 had just come out this time 7 years ago. It doesn’t seem implausible to me that over the next 7 years, we’ll find some way for models to learn on the job.
订阅以获取未来的文章和博客。
You might react, “Wait you made this huge fuss about continual learning being such a handicap. But then your timeline is that we’re 7 years away from what would at minimum be a broadly deployed intelligence explosion.” And yeah, you’re right. I’m forecasting a pretty wild world within a relatively short amount of time.
AGI timelines are very lognormal. It's either this decade or bust. (Not really bust, more like lower marginal probability per year - but that’s less catchy).AI progress over the last decade has been driven by scaling training compute of frontier systems (over 4x a year). This cannot continue beyond this decade, whether you look at chips, power, even fraction of raw GDP used on training. After 2030, AI progress has to mostly come from algorithmic progress. But even there the low hanging fruit will be plucked (at least under the deep learning paradigm). So the yearly probability of AGI craters.
[](https://substackcdn.com/image/fetch/$s_!cAOq!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2a0dc7f9-1224-47b3-9508-d56bfd9fe14f_1029x562.jpeg)
This means that if we end up on the longer side of my 50/50 bets, we might well be looking at a relatively normal world up till the 2030s or even the 2040s. But in all the other worlds, even if we stay sober about the current limitations of AI, we have to expect some truly crazy outcomes.
Subscribe for future posts and blog posts.
[](https://substack.com/profile/11643651-sharmake-farah)[](https://substack.com/profile/4695131-guildmaster-brendon)[](https://substack.com/profile/757328-vincent-weisser)[](https://substack.com/profile/2224934-mehtap-ozkan)[](https://substack.com/profile/110637722-ozgur)
[](https://substack.com/profile/1831134-daniel-kokotajlo?utm_source=comment)
好文章!这基本上也是我的思考方式。那么我们的时间线为何不同呢?
Great post! This is basically how I think about things as well. So why the difference in our timelines then?
——嗯,实际上它们并没有那么不同。我的智能爆炸中位数现在是 2028 年(比写《AI 2027》时多了一年),这意味着《AI 2027》中描述的超级程序员里程碑大约在 2028 年初,我认为这大致对应你描述的“能端到端处理税务”的里程碑,你给这个里程碑在 2028 年底的概率是 50%。也许这有点粗略;可能更像是月级视野而非周级。但按照我们看到的以及我预期的视野长度增长率,这不到一年……
--Well, actually, they aren't that different. My median for the intelligence explosion is 2028 now (one year longer than it was when writing AI 2027), which means early 2028 or so for the superhuman coder milestone described in AI 2027, which I'd think roughly corresponds to the "can do taxes end-to-end" milestone you describe as happening by end of 2028 with 50% probability. Maybe that's a little too rough; maybe it's more like month-long horizons instead of week-long. But at the growth rates in horizon lengths that we are seeing and that I'm expecting, that's less than a year...
——所以基本上我们唯一严重的分歧在于持续/在线学习,你说 2032 年概率 50%,而我认为是 2028 年底概率 50%。这里我的论点很简单:我认为一旦达到超级程序员里程碑,算法进步的速度会加快,然后你会达到完全的 AI 研发自动化,这又会进一步加速,等等。基本上我认为在那个时期进步会比正常快得多,所以像灵活的在线学习这样直觉上可能要到 2032 年才会出现的创新,反而会在同一年晚些时候出现。
--So basically it seems like our only serious disagreement is the continual/online learning thing, which you say 50% by 2032 on whereas I'm at 50% by end of 2028. Here, my argument is simple: I think that once you get to the superhuman coder milestone, the pace of algorithmic progress will accelerate, and then you'll reach full AI R&D automation and it'll accelerate further, etc. Basically I think that progress will be much faster than normal around that time, and so innovations like flexible online learning that feel intuitively like they might come in 2032 will instead come later that same year.
(作为参考,《AI 2027》描绘了从今天到完全在线学习的渐进过渡,中间阶段看起来像是“每周,然后最终每天,他们堆叠另一个微调运行在额外数据上,包括越来越多的在职真实世界数据。”一个粗糙的、无原则的解决方案在 2027 年初出现,并在年中让位于更优雅、更有效的东西。)
(For reference AI 2027 depicts a gradual transition from today to fully online learning, where the intermediate stages look something like "Every week, and then eventually every day, they stack on another fine-tuning run on additional data, including an increasingly high amount of on-the-job real world data." A janky unprincipled solution in early 2027 that gives way to more elegant and effective things midway through the year.)
[](https://substack.com/profile/106343627-ryan-greenblatt?utm_source=comment)
我同意这篇文章的大部分内容。我也大致认为事情变得疯狂的中位数是 2032 年,我同意在职学习非常有用,而且我也怀疑在没有进一步 AI 进步的情况下会出现大规模的白领自动化。
I agree with much of this post. I also have roughly 2032 medians to things going crazy, I agree learning on the job is very useful, and I'm also skeptical we'd see massive white collar automation without further AI progress.
然而,我认为 Dwarkesh 错误地暗示基于强化学习的微调在本质上不能与人类学习方式相似。
However, I think Dwarkesh is wrong to suggest that RL fine-tuning can't be qualitatively similar to how humans learn.
在文章中,他讨论了 AI 基于人类反馈为自己构建可验证的强化学习环境,然后认为这不够灵活和强大,无法工作,但强化学习可以更类似于人类学习的方式使用。
In the post, he discusses AIs constructing verifiable RL environments for themselves based on human feedback and then argues this wouldn't be flexible and powerful enough to work, but RL could be used more similarly to how humans learn.
我最好的猜测是,人类在职学习的方式主要是注意到什么时候事情进展顺利(或不顺利),然后高效地采样更新(他们的大脑做类似于强化学习更新的事情)。在某些情况下,这基于外部反馈(例如来自同事),在某些情况下基于自我验证:这个人只是观察自己行动的结果,然后判断是好是坏。
My best guess is that the way humans learn on the job is mostly by noticing when something went well (or poorly) and then sample efficiently updating (with their brain doing something analogous to an RL update). In some cases, this is based on external feedback (e.g. from a coworker) and in some cases it's based on self-verification: the person just looking at the outcome of their actions and then determining if it went well or poorly.
所以,你可以想象基于外部反馈和自我验证对 AI 进行强化学习。这将是一个像人类学习一样的“有意的、适应性的过程”。为什么这目前比人类学习效果差?
So, you could imagine RL'ing an AI based on both external feedback and self-verification like this. And, this would be a "deliberate, adaptive process" like human learning. Why would this currently work worse than human learning?
当前的 AI 在两个方面比人类差,这使得强化学习对他们来说(在数量上)更差:
Current AIs are worse than humans at two things which makes RL (quantitatively) much worse for them:
1. 稳健的自我验证:正确判断自己何时做得好/差的能力,且这种能力能抵御针对它的优化。
1. Robust self-verification: the ability to correctly determine when you've done something well/poorly in a way which is robust to you optimizing against it.
2. 样本效率:每次更新中学到多少(可能利用诸如确定什么导致事情进展顺利/不顺利之类的东西,人类当然会利用这一点)。这在外部反馈稀疏时尤其重要。
2. Sample efficiency: how much you learn from each update (potentially leveraging stuff like determining what caused things to go well/poorly which humans certainly take advantage of). This is especially important if you have sparse external feedback.
但是,我认为这些更多是数量问题而非质量问题。AI(和强化学习方法)在这两方面都在改进。
But, these are more like quantitative than qualitative issues IMO. AIs (and RL methods) are improving at both of these.
尽管如此,我认为更好的持续学习路径更可能通过建立在上下文学习之上(也许通过类似 neuralese 的东西,尽管这会大大增加不对齐风险……)。
All that said, I think it's very plausible that the route to better continual learning routes more through building on in-context learning (perhaps through something like neuralese, though this would greatly increase misalignment risks...).
- 对于 Dwarkesh 提到的具体播客任务,似乎简单的微调加上一点强化学习就能解决他的问题。所以,由 AI 运行的自动训练循环可能在这里有效。这只是没有作为易用功能部署。
- For the exact podcasting tasks Dwarkesh mentions, it really seems like simple fine-tuning mixed with a bit of RL would solve his problem. So, an automated training loop run by the AI could probably work here. This just isn't deployed as an easy-to-use feature.
- 对于许多(我认为大多数)有用的任务,AI 受限于“在职学习”之外的其他因素。在自主软件工程中,他们无法在 3 小时内匹配人类,并且通常受限于作为糟糕的智能体或普遍愚蠢/困惑。明确地说,对于 Dwarkesh 提到的播客任务,学习似乎是限制因素。
- For many (IMO most) useful tasks, AIs are limited by something other than "learning on the job". At autonomous software engineering, they fail to match humans with 3 hours of time and they are typically limited by being bad agents or by being generally dumb/confused. To be clear, it seems totally plausible that for podcasting tasks Dwarkesh mentions, learning is the limiting factor.
- 相应地,我猜测我们没有看到人们在正常部署中尝试更复杂的基于强化学习的持续学习的原因是,其他地方有更容易摘的果实,并且通常有其他主要障碍。我同意,如果你有像人类一样的样本效率,这会立即产生强大的结果(例如,你会有非常超人的 AI,大概使用 10^26 FLOP),我只是在谈论更渐进的进步。
- Correspondingly, I'd guess the reason that we don't see people trying more complex RL based continual learning in normal deployments is that there is lower hanging fruit elsewhere and typically something else is the main blocker. I agree that if you had human level sample efficiency in learning this would immediately yield strong results (e.g., you'd have very superhuman AIs with 10^26 FLOP presumably), I'm just making a claim about more incremental progress.
- 我认为 Dwarkesh 对“智能”一词的使用有些非典型,当他说“人类如此有用的原因主要不是他们的原始智力。而是他们建立上下文、审视自己的失败、并在练习任务时积累小改进和效率的能力。”我认为人们通常将一个人在职学习的速度视为智能的一个方面。我同意短期反馈回路智能(如智商测试)和长期反馈回路智能之间存在差异,并且在人类中它们相当相关(而 AI 在长期反馈回路智能方面相对较差)。
- I think Dwarkesh uses the term "intelligence" somewhat atypically when he says "The reason humans are so useful is not mainly their raw intelligence. It's their ability to build up context, interrogate their own failures, and pick up small improvements and efficiencies as they practice a task." I think people often consider how fast someone learns on the job as one aspect of intelligence. I agree there is a difference between short feedback loop intelligence (e.g. IQ tests) and long feedback loop intelligence and they are quite correlated in humans (while AIs tend to be relatively worse at long feedback loop intelligence).
- Dwarkesh 指出:“一个能够在线学习的 AI 可能会很快在功能上成为超级智能,即使在那之后没有算法进步。”这似乎合理,但值得注意的是,如果样本高效学习非常计算昂贵,那么这可能不会那么快发生。
- Dwarkesh notes "An AI that is capable of online learning might functionally become a superintelligence quite rapidly, even if there's no algorithmic progress after that point." This seems reasonable, but it's worth noting that if sample efficient learning is very compute expensive, then this might not happen so rapidly.
- 我认为 AI 可能会通过一系列技巧克服样本效率低的问题,以达到非常高的性能水平(例如,构建大量强化学习环境,在反馈稀缺时使用大量算力学习,由于“一次学习多次部署”策略而从比人类多得多的数据中学习)。我认为我们可能会在匹配顶级人类在职学习样本效率之前看到完全自动化的 AI 研发。值得注意的是,如果你确实匹配了顶级人类学习样本效率(同时使用与人类大脑相似数量的算力),那么我们已经拥有足够的算力,这基本上会立即导致远超人类的 AI(人类一生算力大约是 3e23 FLOP,我们很快将进行 1e27 FLOP 的训练运行)。所以,要么样本效率必须更差,要么至少不可能在不花费更多算力每数据点/轨迹/回合的情况下匹配人类样本效率。
- I think AIs will likely overcome poor sample efficiency to achieve a very high level of performance using a bunch of tricks (e.g. constructing a bunch of RL environments, using a ton of compute to learn when feedback is scarce, learning from much more data than humans due to "learn once deploy many" style strategies). I think we'll probably see fully automated AI R&D prior to matching top human sample efficiency at learning on the job. Notably, if you do match top human sample efficiency at learning (while still using a similar amount of compute to the human brain), then we already have enough compute for this to basically immediately result in vastly superhuman AIs (human lifetime compute is maybe 3e23 FLOP and we'll soon be doing 1e27 FLOP training runs). So, either sample efficiency must be worse or at least it must not be possible to match human sample efficiency without spending more compute per data-point/trajectory/episode.
(我最初在推特上发布了这个(https://x.com/RyanPGreenblatt/status/1929757554919592008),但认为放在这里也可能有用。)
(I originally posted this on twitter (https://x.com/RyanPGreenblatt/status/1929757554919592008), but thought it might be useful to put here too.)