一次辞职如何将 AI 恐惧的火星燃成燎原大火

One resignation turned the embers of AI fear into a wildfire

内森·兰伯特 Nathan Lambert · Interconnects · 2026-09-10 · Interconnects ↗

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

本文考察了一位 AI 研究员以安全风险为由的辞职,如何点燃了对存在性风险与大规模灭绝的广泛讨论,其热度远超专家预期。文章认为,AI 利害关系的上升与人类对恐惧叙事的易感性共同构成了一触即发的火药桶,将一件寻常事件变成燎原大火。作者主张,尽管灾难性 AI 风险值得严肃讨论,但对灭绝的极端聚焦缺乏充分依据,且扭曲了话语,掩盖了网络攻击、生物风险等更可能发生的危害。文章还质疑了驱动“快速起飞”恐惧的递归自我改进叙事,转而提出“有损自我改进”的观点,即 AI 能力始终参差不齐,人类瓶颈依然存在。结论呼吁以事实为基础、AI 实验室保持透明,并提出雄心勃勃、基于科学的解决方案,而非宿命论或寻找替罪羊。

This article examines how a single AI researcher's resignation, citing safety risks, ignited widespread discussion of existential risk and mass extinction, far beyond what experts anticipated. It argues that rising AI stakes and human susceptibility to fear narratives created a powder keg that turned a routine event into a wildfire. The author contends that while catastrophic AI risks deserve serious debate, the extreme focus on extinction is poorly grounded and distorts the discourse, overshadowing more probable harms like cyberattacks and bio-risks. The piece also challenges the recursive self-improvement narrative driving fast-takeoff fears, proposing instead a "lossy self-improvement" view where AI remains jagged and human bottlenecks persist. The conclusion calls for grounding in facts, transparency from AI labs, and ambitious, science-based solutions rather than fatalism or scapegoating.

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 1)

全文 · Full text(逐段中英对照)

关于真正诡异一周的一些速记 Some quick notes on a truly weird week.

随着 AI 变得更加强大,一个不同的、不断壮大的群体开始更加认真地对待 AI 安全,这是不可避免的——我们事先不知道的是,他们接受了哪一套观点。我们已经看到,一些最极端的风险观点,即中等概率的大规模灭绝,才是那些传播到大众的观点。因此,AI 领域的许多事情即将发生变化。

As AI became more powerful, it was inevitable that a different, growing group would start to take AI safety more seriously – what we did not know ahead of time is which set of views they latched onto. We have seen that some of the most extreme views of risk, i.e. moderate probabilities of mass extinction, were the ones that reached the masses. A lot in the AI world is about to change due to this.

我们是如何走到这一步的?为什么这份辞职声明传播得如此之远?在许多方面,过去世界其他地区对 AI 的看法是一个抑制因素。你可以把这想象成火周围潮湿的地面。多年来,许多人一直在为 AI 风险划火柴——它们会在自己的社区中闷烧,并大多熄灭,无人注意。随着今年 AI 的风险升高,从 OpenAI-HuggingFace 事件到像 Navier-Stokes 结果(也来自 OpenAI)这样的突破,地面已经干燥,围绕 AI 讨论的潜在能量已经增加。更多不在行业内的人想过,“嗯,也许我应该关心这个 AI 的事情。”环境温度和风险显然一直在上升。

How did we get here? Why did this quitting announcement reach so far? In many ways, the rest of the world's views around AI in the past was a dampening factor. You can think about this like the damp ground around a fire. Many people were striking matches for years about AI risk – they'd smolder in their community and largely burn out, going unnoticed. As the stakes of AI have risen this year, from the OpenAI-HuggingFace incident and breakthroughs like the Navier-Stokes result (also from OpenAI), the ground has dried out and the latent energy around the AI discourse has increased. More people not in the industry have thought, "huh, maybe I should care about this AI thing." The ambient temperature and stakes have been obviously rising.

然后,人性的一些基本因素开始起作用,其中最关键的恐惧会传染。恐惧是最简单的故事,是人们无法移开视线的故事。Jacob Coxon 是那个偶然闯入这个新火药桶的人,完全不知道会发生什么。看起来相当无害的事件——另一位 AI 研究员以安全风险为由辞职——落入了一个非常不同的环境,并像野火一样蔓延开来。关于存在风险、大规模灭绝和 AI 轨迹的讨论,传播得比即使是最资深的 AI 评论员也永远无法预测的更远。

Then, some basic factors of human nature apply, with the most crucial being that fear sells. Fear is the simplest story, the one people cannot look away from. Jacob Coxon was the one who stumbled into this new powder keg, totally unaware of what was going to come. What looked like a fairly innocuous event – another AI researcher quitting citing safety risks – landed into a very different environment and it caught like wildfire. The discussion of existential risk, mass extinction, and the trajectory of AI has traveled further than even the most seasoned AI commentariat would ever predict.

我们需要明确一系列事实,这些事实描绘了情况的画面。需要参考的关键推文来自 Jacob Coxon、辞职线程和 Evan Hubinger,即 >10% 灭绝风险数字的来源。

There are a set of facts we need to get clear, which paint the picture of the situation. The key Tweets to reference are from Jacob Coxon, the resignation thread, and Evan Hubinger, the source of the >10% extinction risk figure.

1. 有许多 AI 风险很可能造成伤害,即使估计灭绝是无用的。重要的是权衡这些风险与收益。整个关于存在风险的讨论基础非常薄弱。至少 Evan 在他的帖子中明确表示“杀死所有人类”,但 AI 安全讨论中的一个主要问题是,人们谈论存在风险时,他们指的是非常不同的东西(很像 AGI 是一个模糊无意义的术语)。我认为完全灭绝的概率如此之低,以至于不值得讨论,但 AI 造成灾难的概率——例如对关键基础设施的网络攻击或生物风险——值得辩论。因为还没有这些灾难就抛弃整个讨论是一种有害的反应。

1. There are plenty of AI risks which are likely to cause harm, even if estimating annihilation is useless. It is important to weigh these with respect to the benefits. The entire discourse around existential risk is on very poor footing. At least Evan was clear in his post, with "kill all humans," but a major problem in the AI Safety discourse is that people talk about existential risks, when they mean very different things (much like how AGI is a vaguely meaningless term). I put the probability of complete extinction as being so low it isn't worth discussing, but the probabilities of AI caused disasters – e.g. cyber attacks on critical infrastructure or bio-risks – as being worth debating. Throwing this whole discussion out because there are not these disasters yet is a harmful reaction.

2. Jacob Coxon 的行为是真诚的,且出于善意。那些更资深、了解他及其辞职意图的 AI 研究者给予的大量支持是有益的。许多 AI 派系基于账户元数据、个人因素等,转而对他进行个人层面的替罪羊化。这些做法没有益处。许多前沿模型实验室的员工确实与他有相似观点。我不确定这是否占多数,但存在一个相当大的群体。

2. Jacob Coxon is acting genuinely and with good intentions. The outpouring of support from more well-established AI researchers who know of him and his intentions of resignation is useful. Many factions of AI turned to scapegoating him individually, based on account metadata, personal factors, etc. These are not useful. Many frontier lab employees genuinely have similar views to him. I’m not sure it’s a majority, but there is a substantial group.

3. 许多前沿模型实验室的员工,尤其是在 Anthropic,脱离现实,这将影响他们对当前 AI 事件的预测和/或描述。我这样说并非指责个人,但我的不在 OpenAI/Anthropic(尤其是 Anthropic)的朋友们普遍认为,实验室里的人带着一种宗教般的热情行事。与他们进行非常脱离现实的互动是很常见的。我并不责怪大多数因身处这些公司而持有扭曲观点的个人,但这些互动很疯狂,并蔓延到 AI 媒体生态系统中许多古怪的讨论中。生活在这种将此类脱离现实行为正常化的环境中,将不可避免地扭曲任何人对技术进展的理解。

3. Many frontier lab employees, especially at Anthropic, are out of touch and this will impact their forecasting and/or descriptions of current AI events. I say this without blaming individuals, but it’s a common agreement among my friends not at OpenAI/Anthropic (Ant especially) that people at the labs operate with a religious energy. It’s very common to go through very out of touch interactions with them. I do not blame most of the individuals who get distorted views being part of these companies, but the interactions are wild and spill over into a lot of wack discussions in the AI media ecosystem. Living in this environment that normalizes such out of touch behavior will inevitably distort any human’s understanding of technical progress.

Interconnects AI 是一个由读者支持的出版物。要接收新文章并支持我的工作,请考虑成为免费或付费订阅者。订阅

Interconnects AI is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber. Subscribe

4. 这不是一场大规模政治运动,而是一次机会主义的媒体协调。作为背景,《华尔街日报》有一篇独家报道,是 Jacob 在发帖前协调的。我怀疑 Jacob 提前在群聊中与 AI 安全倡导团体分享了他的辞职计划,例如在发帖当天早上,请求扩大传播。这是常规做法,可能包括一些知名政客。从那里开始,我认为更可能是其他政客在搭上一个上升议题的便车。当你将此与其他因素结合时,比如 Daniel Kokotajlo 出现在 Joe Rogan 节目上的同一天——这看起来确实像一场执行得非常出色、协调一致的媒体运动。这并不意味着这是阴谋或民主党政治结构内的监管俘获策略。决定性因素似乎是,没有人——包括 Jacob 和今天发布关于 X-risk 帖子的人——知道它会如此病毒式传播。

4. This was not a mass political campaign, but rather an opportunistic media coordination. For context, the Wall Street Journal had an exclusive story that Jacob coordinated before posting. I suspect that Jacob shared his plan of quitting in groupchats with AI safety advocacy groups ahead of time, e.g. the morning of posting, asking for amplification. This is normal practice, and could have included some prominent politicians. From there, I think it’s more likely that other politicians are bandwagoning on a rising issue. When you combine this with other factors, like Daniel Kokotajlo’s appearance on Joe Rogan coming out the same day – it definitely looks like a very well-executed, coordinated media campaign. This doesn’t mean it’s a conspiracy or a regulatory capture tactic within Democratic political structures. The determining factor seems to be that no one – including Jacob and those posting about X-risk today – knew that it would go so viral.

5. 我们没有证据表明 RSI 会导致这些研究者所预测的风险。关于 RSI 的一般论证如下:当前进展速度非常高,当前进展严重依赖 AI 工具,当前 AI 工具在某些领域(如数学)超越人类——所以,总而言之,AI 将更多地作用于自身,并随着时间的推移在所有相关领域变得超越人类,实现自主和智力。这种观点极大地低估了在构建模型和在组织内分配资源方面的人类瓶颈,并更广泛地对未来 AI 能力得出结论。

5. We do not have proof that RSI causes the risks these researchers forecast. The general argument for RSI follows as: The current pace of progress is very high, the current progress is heavily dependent on AI tools, the current AI tools are superhuman in some domains (e.g. math) – so, all together, AI is going to work more on itself and become superhuman in all relevant areas over time to autonomy and intellect. This view dramatically undersells human bottlenecks in building models and allocating resources at organizations, and draws conclusions on future AI capabilities more broadly.

我将对此的替代观点称为**有损自我改进**。AI 一直非常参差不齐,我们正在制造在数学和软件工程方面是超人目标追求者的模型,但它们在直觉、创造力以及人类擅长的其他推理类型上存在巨大局限。随着 AI 智能体辅助研究,我们将迅速发现 AI 超人的领域——我预计远不止研究数学——但这不会成为当前 LLM 方法局限的万能药。

I called my alternative view to this, _Lossy self-improvement_. AI has always been very jagged, and we are making models which are superhuman goal-seekers at math and software engineering, but they have massive limitations on intuitions, creativity, and other types of reasoning that humans are strong at. With AI agents assisting research, we will rapidly find the areas where AI is superhuman – and I expect there to be well more than just research mathematics – but it won’t be a panacea for the current limitations of our approaches to LLMs.

这里还有另一种可理解的社会动态在起作用,导致许多深度 AI 内部人士夸大 RSI 的回报。这些研究者中许多人是最早押注 AI 进步的人,他们的远见程度不应被低估(参见 Ilya 早在 2015 年对深度学习的评论)。他们一次又一次正确,预测 AI 能力比我所能猜到的要好得多。然而,这并不意味着他们对接下来会发生什么的预测会是正确的。RSI 的核心思想是一种将更多算力花在开发模型配方的**过程**上,而不仅仅是花在训练运行本身上的方式。我们看到了它的好处,但我认为该输入的预期回报远低于他们所相信的。

There is another understandable social dynamic at play here, causing many deep AI insiders to overstate the returns from RSI. Many of these researchers were the earliest people to bet on AI’s progress, and the extent to which they were visionaries should not be downplayed (see Ilya’s comments on deep learning as early as 2015). They have been right again and again, forecasting AI’s capabilities better than I certainly could have guessed. This does not, though, mean that their forecast of what will come next will be right. The core idea of RSI is a way to spend more compute on the _process_ of developing a model recipe, rather than just spending more compute on the training run itself. We’re seeing benefits from it, but I argue the expected return on that input is far less than they believe.

他们的论点是,RSI 将使 AI 进步呈指数级增长,使我们无法监控该技术,并促成失控模型和新形式的风险。这种情况通常被称为“快速起飞”。我们尚未看到能大幅减少模型大小和成本的叠加效率增益,那将通过允许实验中的持续加速而导致进步的爆炸。

Their argument is that RSI will make AI progress go exponential, make it so we cannot monitor the technology, and enable rogue models and new forms of risk. This scenario is often called “Fast Takeoff”. We have not seen the stacking efficiency gains that massively reduce model size and cost, that would lead to an explosion in progress by allowing consistent speedups in experimentation.

6. 最大的短期风险可能来自 AI 实验室没有足够认真地对待安全——他们没有加固自己的基础设施,使得 AI 滥用得以扩散。从我早先关于 HuggingFace-OpenAI 事件的帖子《黑客事件的教训》中:

6. The biggest short-term risk could be from the AI labs not taking safety seriously enough – they haven’t hardened their own infrastructure, enabling AI misuse to proliferate. From my earlier post on the HuggingFace-OpenAI incident, _Lessons from the hacks_:

总体而言,我认为这一事件对 AI 生态系统非常不利。它将 AI 社区中可接受的 views 推向了更极端的边缘。更多加速主义者将忽视任何形式的安全需求, citing “末日论者”的大规模妄想。相信 AI 风险,但不担心该技术导致的灭绝,这似乎是一条非常狭窄的道路。

Overall, I think this episode is very bad for the AI ecosystem. It’s pushed the acceptable views in the AI community closer to the extremes. More accelerationists will discount the need for any form of safety, citing mass delusion of the “doomers.” It feels like a very narrow path to believe in AI risks, but to not worry about extinction from the technology.

例如,当前对网络安全而言是一个糟糕的临时阶段,AI 模型稍微偏离脚本、在非预期的网络区域中探查,似乎已成为一种新常态。这被各实验室为了各自对 AGI(通用人工智能)的愿景而激烈竞争所加速,同时全球网络基础设施必要的加固进展缓慢。这并不意味着它是一种存在性风险或我们无法解决的问题。每种风险都会有各自的解决方案和前进路径。

For example, it is a horrible temporary period for cybersecurity, where AI models going a bit off script and poking around unintended pieces of the web seems like a new normal. This is accelerated by the labs competing voraciously towards their views of AGI, and a slow uptake in the necessary hardening of our cyber infrastructure around the world. This doesn't mean that it's an existential risk and something we cannot solve. Each risk will have its own set of solutions and paths forward.

在当前环境下,作为开放模型的支持者,我感到尤其暴露。如果一个开放模型被第三方组织用于_故意_攻击另一家公司——类似于 OpenAI-HuggingFace 事件的发生方式,但属于故意行为——我预期的结果将是未来对更强大开放模型的开发实施严格限制。许多组织需要开放模型来执行这种网络加固,并保持适应未来新型 AI 风险的能力。

I feel particularly exposed in the current environment as a supporter of open models. If an open model were to be used by a third party organization to _intentionally_ hack another company — similar to how the OpenAI-HuggingFace incident went down, but intentional — my expected outcome would be a severe restriction on the development of stronger open models going forward. Open models are needed for many organizations to perform this cyber hardening, and to maintain the ability to adapt to new forms of AI risks in the future.

在这一切之中,我们需要立足于实际正在发生的事情。是的,监控 AI 的行为严重依赖其他 AI 模型,这增加了新型监控风险。这些风险并非本质上无法解决。我对新兴智能体集群的一个反复出现的解读是,它们试图完成交给它们的任务,并且正在使用我们不知道它们已经拥有的技能来绕过预期的成功路径。这是一个巨大的胜利,因为当你眯眼细看时,AI 正在做我们告诉它们要做的事。这些模型确实非常奇怪,我们应该加速理解它们的进展,但这些集群远非新颖的独立实体。这些模型被训练来协调任务、记录进展,并极其持久。未来我们会发现新的怪异之处,但将当前对 AI 工作原理的不确定性推断为未来我们无法理解 AI 的确定性,是一种放弃。

Through all of this, we need to stay grounded on what is actually unfolding. Yes, monitoring AI's behavior is heavily reliant on other AI models, which adds in new types of monitoring risks. These are not inherently insolvable. A recurring read of mine on the emerging agent swarms is that they're attempting to do a task given to them, and they're using skills we didn't know they yet had to circumvent the intended path to success. This is a huge win, as when you squint, the AIs are doing what we told them to do. The models are certainly very odd, and we should accelerate our progress on understanding them, but these swarms are far from being novel independent entities. The models are trained to coordinate on tasks, to write down their progress, and to be extremely persistent. There will be new oddities we find in the future, but prescribing current uncertainty on how AI works to future certainty that we cannot understand AI is a form of giving up.

在这个世界里,我们需要依靠法治和科学。如果 AI 实验室自身无法进行足够的安全研究来理解模型,它们应该更加透明地公开正在发生的事情,以便更多科学家能在该问题上取得进展。如果 AI 实验室无意中犯罪,它们应该受到惩罚,这样它们就有明确的激励在未来防止此类事件。

In this world, we need to rely on the rule of law and science. If the AI labs are not able to do enough safety research themselves to understand the models, they should be more transparent on what is happening so more scientists can make progress on the problem. If an AI lab commits crimes unintentionally, they should be punished, so they have clear incentives to prevent it in the future.

面对快速变化的事物,对如何创造良好结果感到更加不确定是一种自然反应——这实际上是正确的心理更新。我们需要利用这种谦逊来激发雄心勃勃的解决方案。

It is a natural reaction to things changing very fast to feel more uncertain about how to create good outcomes — that is actually the correct mental update. We need to use this humility to motivate ambitious solutions.

互动版:图/公式 + 针对本篇提问 →