为什么我仍未认同真正的递归自我改进

Why I still haven’t bought into true RSI

内森·兰伯特 Nathan Lambert · Interconnects · 2026-09-19 · Interconnects ↗

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

本文反对真正的递归自我改进(RSI)即将到来的观点,认为当前的人工智能安全焦虑,更多是被前沿实验室的狂热文化以及数千个并发智能体可见的生产力所放大,而非源于真正未公开的突破。作者的基准判断是“有损自我改进”:在规模定律带来的指数级成本下,可自动化的研究范围太窄,不足以大幅加速进展;并行智能体收益递减;资源与政治瓶颈也制约着强大的大语言模型。因此,RSI 更可能让现有模型变得极其便宜和高效,改进推理时扩展与多智能体系统,而不是扩展智能的峰值——后者仍是最难推动的指数。由于大语言模型的智能参差不齐,而人类角色是“人类形状”的,AI 不会离散地跨过诸如远程工作者或 AI 研究者这样的门槛;扩散将是缓慢且长尾的,人类的理解与沟通仍是关键瓶颈。结论是:在出现更强证据之前,有损自我改进应仍是基准轨迹,而加剧的灭绝风险论调是错位的。

The article argues against the imminent arrival of true recursive self-improvement (RSI), contending that current AI safety anxiety is amplified by the frenetic culture of frontier labs and the visible productivity of thousands of concurrent agents rather than by genuine, undisclosed breakthroughs. The author's baseline is "lossy self-improvement": automatable research is too narrow to massively accelerate progress given scaling laws' exponential costs, parallel agents show diminishing returns, and resource and political bottlenecks constrain strong LLMs. RSI is therefore more likely to make existing models vastly cheaper and more efficient, improving inference-time scaling and multi-agent systems, than to expand peak intelligence, which remains the hardest exponential to move. Because LLM intelligence is jagged and human roles are human-shaped, AI will not discretely cross thresholds like remote worker or AI researcher; diffusion will be slow and long-tailed, with human understanding and communication remaining the key bottleneck. The conclusion is that lossy self-improvement should remain the baseline trajectory until stronger evidence emerges, and that heightened extinction-risk discourse is misplaced.

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 1)

全文 · Full text(逐段中英对照)

一位“AI 温和派”对近期事件与前沿模型发展轨迹的看法 An “AI moderate’s” view on recent events and the trajectory of frontier models.

我们正处于这样一个时代:少数组织正在使用数千个并发智能体来改进其流程和产出。这些组织恰好就是前沿 AI 实验室,尤其是 OpenAI 和 Anthropic。在过去几周里,我一直在思考,这些组织中如此多的员工迅速更新对 AI 进展速度及相关风险的预期,意味着什么。

We're in an era where a few organizations are using thousands of concurrent agents to improve their processes and output. These organizations happen to be just the frontier AI labs, in particular OpenAI and Anthropic. In the last few weeks, I've been pondering what it means for so many employees across these organizations to rapidly update their expectations for the pace of AI progress and associated risks.

我的一个核心观点是,前沿实验室以及旧金山 AI 圈更广泛的狂热竞争文化,营造了一种放大任何 AI 担忧的环境。这有一些好处,能让更广泛的受众意识到 AI,因为恐惧好卖,但夸大风险时间线或严重程度会产生负面的二阶效应。我记得 2023 年和 2024 年许多喧嚣的 AI 安全辩论,以及它们对开源 AI 可行性的阴云——当时的主要风险并未在预测的时间线内出现。

A core perspective I have is that the frontier labs and broader frenetic, competitive culture in the San Francisco AI scene set up an environment that amplifies any AI concern. This has some benefits in causing more general audience awareness of AI, as fear sells, but exaggerating risk timelines or severity will have negative second-order effects. I remember many loud AI safety debates, and their associated clouds over the viability of open-source AI, in 2023 and 2024 — the primary risks then did not arrive in the forecasted timelines.

即便在一年前,这两家关键实验室的普通员工就对 AI 风险和进展速度非常焦虑,尤其是当智能体在 2026 年初获得更强的产品市场契合度时。这种文化前提,一旦暴露于数千个智能体将不断在你的业务中相当高效地工作的现实,只会加剧这种焦虑。从这种焦虑,以及像 OpenAI-HuggingFace 这样的事件,到灭绝风险,这一步感觉非常宗教化。

The general populace of these two key labs was very anxious about AI risks and the rate of progress even a year ago, and especially as agents got stronger product-market fit at the start of 2026. This cultural precondition, when exposed to the reality that thousands of agents will constantly be working fairly productively in your business, will only increase this anxiety. The step from this anxiety, and incidents like OpenAI-HuggingFace, to extinction risks feels very religious.

Richard Ngo 对这种情况有一个贴切的总结:

Richard Ngo had an apt summary of the situation:

个人而言,我认为这一观点与我在真正递归自我改进(RSI)的替代情景中概述的内容高度一致,我称之为有损自我改进。这一观点的总结是:

Personally, I think this view aligns closely to what I outlined in my alternate scenario to true recursive self-improvement (RSI), which I called lossy self-improvement. A summary of this view is that:

1. 可自动化的研究范围过于狭窄,面对缩放定律带来的指数级成本,无法实现进展的大规模净加速;

1. Automatable research is too narrow to achieve a massive net acceleration in progress, in the face of scaling laws' exponential costs,

2. 更多 AI 智能体并行所带来的收益递减是真实存在的;以及

2. Diminishing returns of more AI agents in parallel are real, &

3. 资源瓶颈与政治因素是构建强大 LLM 的主要制约(而 AI 对此的加速作用十分有限)。

3. Resource bottlenecks and politics are a major factor in building strong LLMs (and AI can do much less to accelerate this).

因此,我需要在上述因素、文化温度的潜在上升,以及实验室可能已发现尚未公开的真正可怕的具体突破之间进行权衡。我的预期是,当前更多的 AI 安全担忧集中在前者——即规模化智能体的运作——但我对此高度不确定。基础性的、基于想象力的 AI 突破,会让我将 RSI 时间线从“更接近于一种在指数级成本(缩放定律)面前维持进展的工具”,调整为“更不可预测和/或不稳定的东西”。

So, I'm left balancing the above, latent increase in the cultural temperature with the potential that the labs have seen genuinely scary, specific breakthroughs that are not public yet. My expectation is that more of the current AI safety concern is on the former – scaled agents working – but I hold high levels of uncertainty here. Foundational, imagination-based AI breakthroughs are the sort of thing that would make me update my RSI timelines from closer to a tool to sustain progress in the face of exponential costs (scaling laws), to something more unpredictable and/or unstable.

近期关于 RSI 的最佳资源包括 Dwarkesh 与 Noam Brown 的播客,以及与 John Schulman、Beren Millidge 和 Charlie O'Neill 三人的播客。我从这两期播客中获得了若干重要思考。

Some of the best recent resources on RSI have been Dwarkesh's podcasts with Noam Brown and the trio of John Schulman, Beren Millidge and Charlie O'Neill. I have a few important reflections from both of them.

首先,与 Noam Brown 的播客让我深刻认识到,大规模推理能力在短期内能带来多大的加速。这些实验室将投入数千个智能体去解决重要的、可度量的问题。与此同时,它们可用的算力将持续扩张。我怀疑,随着总算力规模上升,实验室能否负担将固定比例的算力用于内部研发,尤其是它们计划 IPO,并面临对基本经济模型的更严格审视。重要的是,不要将推理时扩展(inference-time scaling)的巨大进步——这一动态本应相当可预测——与 RSI 的产出混为一谈,后者具有高度不确定性。

First, the podcast with Noam Brown made me internalize how big of a short-term acceleration mass inference capacity is. These labs will throw thousands of agents at important, measurable problems. At the same time, compute capacity available to them is going to continue to scale. I have my doubts that the labs can afford to spend a constant portion of this compute on internal R&D as the total volume goes up, especially with plans to IPO, as they face increased scrutiny on basic economics. It is important to not confuse massive steps in inference-time scaling, a dynamic which should be fairly predictable, with being the outputs of RSI, which is highly uncertain.

其次,三人播客讨论技术能力现状时,引发了我一种尚未完全消化的意外反应。在播客的前一个小时左右,他们辩论了强化学习(RL)、蒸馏、扩展、推理时算力等的作用,我发现自己强烈认同他们观点的分布。简而言之,我们当前的技术有效,能让我们解决已知如何表述的问题,但在大多数部分可验证的领域中,它们并不能神奇地泛化到未知的、更难的问题(即数学的进步是例外,而非规律)。

Second, the trio podcast debating the state of the art in technical capacities induced more of a surprising reaction that I haven't fully settled. Through the first hour or so of this podcast, where they debate the role of RL, distillation, scaling, inference-time compute, etc., I found myself strongly agreeing with the distribution of claims. A TLDR would be that our current techniques work and let us solve problems we know how to state, but they don't result in a magical level of generalization to unknown, harder problems in most partially verifiable domains (i.e. progress in math is an exception, rather than a rule).

这期播客令人意外的是结尾部分,他们在预测 AI 各种阈值的时间线。我让 GPT-6-Astra 总结了 Dwarkesh 提出的三个问题的回答,问题形式为“AI 何时达到 X 能力”:

The surprise of this podcast was the end, where they were predicting timelines for various thresholds of AI. I had GPT-6-Astra summarize the answers provided to three questions from Dwarkesh, of the form "when will AI reach X ability":

大致来说,讨论 RSI 时一个反复出现的问题是智能缺乏明确界定。智能的锯齿状特性意味着我们需要在具体的、可度量的任务中讨论阈值。大语言模型(LLM)智能的本质与人类 _截然不同_,而我们预测的角色却是人类形态的。因此,AI 不会像远程工作者或 AI 研究员那样离散地跨越这些阈值。这是一个缓慢的扩散过程,长尾形式将始终存在。

Roughly, a recurring problem when discussing RSI is a lack of specification in intelligence. The jaggedness of intelligence means that we need to discuss thresholds in specific, measurable tasks. The nature of LLMs' intelligence is shaped _very differently_ than humans, and the roles we forecast are human-shaped. AIs, therefore, do not cross these thresholds like remote worker or AI researcher discretely. It's a slow diffusion, and a form of long tail will always exist.

以 AI 研究者的生产力为例。许多人低估了科学中与同事沟通和制定标准所占的比重。我确实相信在不久的将来,实验设计与测试的周期会快 10 倍,但不相信假设生成和直觉构建也能如此。加速理解将成为关键瓶颈——尽管所有 AI 工具都得到了大幅改进,但人类在这方面的能力只会略有提升。科学本质的一大改进将是让人类能在此投入更多时间,而非让人类在这方面呈指数级变强。

Take the case of productivity of AI researchers. Many people under-index how much of science is communication and standard setting with colleagues. I do buy the cycle of experiment design and testing being 10x faster in the near future, but not hypothesis generation and intuition building. Accelerating understanding will be the key bottleneck – and it is one that despite all of the AI tools getting massively improved, humans will only improve marginally in their capability. A big improvement in the nature of science will be enabling humans to invest more time here, not them becoming exponentially better at it.

这可以联系到 Noam 的播客。在不久的将来,智能体集群将能有效解决那些答案可验证的、清晰开放的问题。同样地,在改进 AI 模型方面,RSI 在提升效率上远比扩展智能峰值更有帮助。这是因为 LLM 服务有明确的可测量、可塑的指标可供改进。这将带来更好的推理时扩展和更高效的多智能体系统。

This links back to the Noam podcast. Agent swarms in the near future will be effective at solving clear, open problems with verifiable answers. In this vein, when it comes to improving AI models, RSI is much more helpful at efficiency rather than expanding peak intelligence. This is due to the fact that LLM serving has clear metrics you want to improve that are measurable and malleable. This’ll enable better inference-time scaling and more efficient multi-agent systems.

尽管如此,我始终无法忽视一个事实:我们所有的缩放定律都表明,要让智能获得线性提升,就需要指数级的算力和资源。RSI 有望让现代 LLM 变得极其便宜。那些显示 LLM 在给定智能水平下成本呈指数级下降的趋势很可能会加速。对实验室而言,一个关键因素将是**提高**利润率,因为在固定智能水平下若价格竞争激烈,收入可能面临下行压力——杰文斯悖论很可能占上风,从而造就强大的企业。

Still, I cannot get past the fact that all of our scaling laws show that you need exponential compute and resources to make linear improvements in intelligence. RSI is poised to make modern LLMs vastly cheaper. Trends that have shown LLMs get exponentially cheaper at a given intelligence are likely to accelerate. A crucial factor for the labs will be _increasing_ margins as revenue could potentially have negative pressure if there’s fierce competition in lowering prices at a fixed intelligence level — Jevons paradox will likely prevail, resulting in strong businesses.

RSI 因素在改进 LLM 拼图中诸如管理复杂的后训练配方等部分时,会困难得多。John Schulman 有几段关于实验室后训练现状的言论,我深表赞同:

RSI factors will have a much harder time improving pieces of the LLM puzzle like managing complex post-training recipes. There were a few quotes from John Schulman that I strongly agree with on the state of post-training at the labs:

这些任务对当前的 LLM 来说格外困难。是的,随着业界仍在快速扩展与这些领域相关的强化学习环境,它们会变得更好,但这种范式不会永远持续。在不久的将来,构思、构建和测试能真正挑战领先 LLM 的新环境可能会变得指数级困难——这些**困难**环境正是强化学习中作为学习信号至关重要的部分。

These tasks are uniquely hard for current LLMs. Yes, they’ll get better as the industry is still rapidly scaling RL environments related to these domains, but this paradigm does not last forever. In the near future, it could become exponentially harder to conceive, build, and test new environments that meaningfully challenge the leading LLMs – these _hard_ environments are the ones that are crucial as a learning signal in RL.

OpenAI 和 Anthropic 已经分享了大量与 RSI 相关的内部测量数据,而我目前的解读是,实验室内部自动化程度提升最大的领域是软件工程、监控日志、管理计划实验以及其他相当常规(但并不总是简单)的任务。例如,最近 Claude Fable 5.1 与 Mythos 5.1 系统卡中的这段表述让我感到意外:

OpenAI and Anthropic have shared a good amount of internal measurements related to RSI, and my current read is that the biggest takeoff in automation within the labs is in tasks like software engineering, monitoring logs, managing planned experiments, and other fairly routine (but not always easy) tasks. For example, I was surprised by this language in the recent Claude Fable 5.1 & Mythos 5.1 System Card:

总而言之,我认为我们正在_对抗_的最难指数增长在于峰值智能。这是最难推动、甚至最难加速的一项。尽管如此,对于 RSI 的极早期阶段,我的心智模型更多是将推理时算力大规模扩展并扩散到 AI 研究及相关活动中,这方面存在大量唾手可得的成果。仅此一点,就仍有望带来经济上的变革性影响。它还可能释放更多资源来推动 AI 扩散,而 AI 扩散是释放 AI 大部分潜在益处的关键瓶颈。

Altogether, I think the hardest exponential we are _fighting_ is on peak intelligence. That is the hardest one to budge or even accelerate. Still, my mental model for the very early innings of RSI is more of massively scaling and diffusing inference-time compute to AI research and related activities, which has a large amount of low-hanging fruit available. This, on its own, is still poised to be economically transformative. It may also unlock more resources to push on AI diffusion, which is the crucial bottleneck in unlocking much of the potential benefits of AI.

就目前而言,在出现更多证据之前,有损自我改进仍是我对进展轨迹的基本判断,而关于灭绝风险的讨论增多似乎非常不合时宜。一如既往,AI 领域的情况可能变化很快。

For now and until more evidence emerges, lossy self-improvement remains my baseline on the trajectory of progress, and the increased discussion of extinction risk seems very misplaced. As always, things can change fast in AI.

互动版:图/公式 + 针对本篇提问 →