达里奥·阿莫迪——技术的青春期

The Adolescence of Technology

达里奥·阿莫迪 Dario Amodei · Anthropic · 2026-01-26 · Dario Amodei Essay ↗

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

2. 一种令人惊讶且可怕的赋能

* 2. A surprising and terrible empowerment

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 15)

全文 · Full text(逐段中英对照)

目录 Contents

* 2. 一种令人惊讶且可怕的赋权

* 2. A surprising and terrible empowerment

技术的青春期 The Adolescence of Technology

面对并克服强大 AI 的风险

Confronting and Overcoming the Risks of Powerful AI

在卡尔·萨根著作《接触》的电影版中有一个场景:主角,一位探测到来自外星文明第一个无线电信号的天文学家,正在被考虑作为人类代表去会见外星人。国际小组面试她时问道:“如果你只能问(外星人)一个问题,你会问什么?”她回答:“我会问他们,‘你们是怎么做到的?你们是如何进化,如何度过这个技术青春期而没有自我毁灭的?’”当我思考人类目前在 AI 方面的处境——我们正处于什么边缘时——我的思绪不断回到那个场景,因为这个问题对我们当前的处境非常贴切,我希望我们有外星人的答案来指引我们。我相信我们正在进入一个成年礼,既动荡又不可避免,这将考验我们作为一个物种的本质。人类即将获得几乎难以想象的权力,而我们的社会、政治和技术系统是否拥有成熟度来运用它,这一点非常不确定。

There is a scene in the movie version of Carl Sagan’s book _Contact_ where the main character, an astronomer who has detected the first radio signal from an alien civilization, is being considered for the role of humanity’s representative to meet the aliens. The international panel interviewing her asks, “If you could ask [the aliens] just one question, what would it be?” Her reply is: “I’d ask them, ‘How did you do it? How did you evolve, how did you survive this technological adolescence without destroying yourself?” When I think about where humanity is now with AI—about what we’re on the cusp of—my mind keeps going back to that scene, because the question is so apt for our current situation, and I wish we had the aliens’ answer to guide us. I believe we are entering a rite of passage, both turbulent and inevitable, which will test who we are as a species. Humanity is about to be handed almost unimaginable power, and it is deeply unclear whether our social, political, and technological systems possess the maturity to wield it.

在我的文章《优雅的机器》中,我试图描绘一个文明已经步入成年的梦想,其中风险已被解决,强大的 AI 被熟练且富有同情心地应用于提高每个人的生活质量。我提出 AI 可以为生物学、神经科学、经济发展、全球和平以及工作和意义带来巨大进步。我认为给人们一些鼓舞人心的东西去奋斗是很重要的,而 AI 加速主义者和 AI 安全倡导者似乎——奇怪地——都未能做到这一点。但在当前这篇文章中,我想直面这个成年礼本身:描绘我们将要面临的风险,并尝试开始制定一个战胜它们的作战计划。我深信我们有能力获胜,相信人类的精神和崇高,但我们必须正视形势,不抱幻想。

In my essay _Machines of Loving Grace_, I tried to lay out the dream of a civilization that had made it through to adulthood, where the risks had been addressed and powerful AI was applied with skill and compassion to raise the quality of life for everyone. I suggested that AI could contribute to enormous advances in biology, neuroscience, economic development, global peace, and work and meaning. I felt it was important to give people something inspiring to fight for, a task at which both AI accelerationists and AI safety advocates seemed—oddly—to have failed. But in this current essay, I want to confront the rite of passage itself: to map out the risks that we are about to face and try to begin making a battle plan to defeat them. I believe deeply in our ability to prevail, in humanity’s spirit and its nobility, but we must face the situation squarely and without illusions.

与讨论益处一样,我认为以谨慎和深思熟虑的方式讨论风险也很重要。特别是,我认为关键是要:

As with talking about the benefits, I think it is important to discuss risks in a careful and well-considered manner. In particular, I think it is critical to:

* 避免末日论。在这里,我所说的“末日论”不仅仅是指认为末日不可避免(这是一种既错误又自我实现的信念),更广泛地说,是以一种准宗教的方式思考 AI 风险。11 这与我在《优雅的机器》中提出的观点对称,我在那篇文章开头说,AI 的好处不应被视为救赎的预言,重要的是要具体和脚踏实地,避免夸大其词。最终,救赎的预言和末日的预言对于面对现实世界都是无益的,原因基本相同。许多人多年来一直以分析和清醒的方式思考 AI 风险,但我的印象是,在 2023-2024 年对 AI 风险担忧的高峰期,一些最不理智的声音占据了主导地位,通常通过耸人听闻的社交媒体账号。这些声音使用令人反感的语言,让人联想到宗教或科幻小说,并呼吁采取极端行动,却没有提供能够证明其合理性的证据。当时就很清楚,反弹是不可避免的,这个问题会变得文化上两极分化,从而陷入僵局。22 Anthropic 的目标是在这种变化中保持一致。当谈论 AI 风险在政治上受欢迎时,Anthropic 谨慎地倡导对这些风险采取明智和基于证据的方法。现在谈论 AI 风险在政治上不受欢迎,Anthropic 继续谨慎地倡导对这些风险采取明智和基于证据的方法。 截至 2025-2026 年,钟摆已经摆动,AI 机遇而非 AI 风险正在推动许多政治决策。这种摇摆是不幸的,因为技术本身并不关心什么是时尚,而我们在 2026 年比 2023 年更接近真正的危险。教训是,我们需要以现实、务实的方式讨论和应对风险:冷静、基于事实,并能够经受住潮流的变迁。

* Avoid doomerism.Here,I mean “doomerism” not just in the sense of believing doom is inevitable (which is both a false and self-fulfilling belief), but more generally, thinking about AI risks in a quasi-religious way.11 This is symmetric to a point I made in _Machines of Loving Grace_, where I started by saying that AI’s upsides shouldn’t be thought of in terms of a prophecy of salvation, and that it’s important to be concrete and grounded and to avoid grandiosity. Ultimately, prophecies of salvation and prophecies of doom are unhelpful for confronting the real world, for basically the same reasons. Many people have been thinking in an analytic and sober way about AI risks for many years, but it’s my impression that during the peak of worries about AI risk in 2023–2024, some of the least sensible voices rose to the top, often through sensationalistic social media accounts. These voices used off-putting language reminiscent of religion or science fiction, and called for extreme actions without having the evidence that would justify them. It was clear even then that a backlash was inevitable, and that the issue would become culturally polarized and therefore gridlocked.22 Anthropic’s goal is to remain consistent through such changes. When talking about AI risks was politically popular, Anthropic cautiously advocated for a judicious and evidence-based approach to these risks. Now that talking about AI risks is politically unpopular, Anthropic continues to cautiously advocate for a judicious and evidence-based approach to these risks. As of 2025–2026, the pendulum has swung, and AI opportunity, not AI risk, is driving many political decisions. This vacillation is unfortunate, as the technology itself doesn’t care about what is fashionable, and we are considerably closer to real danger in 2026 than we were in 2023. The lesson is that we need to discuss and address risks in a realistic, pragmatic manner: sober, fact-based, and well equipped to survive changing tides.

* 承认不确定性。我在这篇文章中提出的担忧有很多种方式可能变得无关紧要。这里没有任何内容旨在传达确定性甚至可能性。最明显的是,AI 可能根本不会像我设想的那样快速发展。33 随着时间的推移,我对 AI 的发展轨迹以及它将在各个领域超越人类能力的可能性越来越有信心,但仍存在一些不确定性。 或者,即使它确实快速发展,这里讨论的部分或全部风险可能不会出现(那将很好),或者可能存在我未考虑到的其他风险。没有人能完全自信地预测未来——但无论如何,我们必须尽力规划。

* Acknowledge uncertainty.There are plenty of ways in which the concerns I’m raising in this piece could be moot. Nothing here is intended to communicate certainty or even likelihood. Most obviously, AI may simply not advance anywhere near as fast as I imagine.33 Over time, I have gained increasing confidence in the trajectory of AI and the likelihood that it will surpass human ability across the board, but some uncertainty still remains. Or, even if it does advance quickly, some or all of the risks discussed here may not materialize (which would be great), or there may be other risks I haven’t considered. No one can predict the future with complete confidence—but we have to do the best we can to plan anyway.

* 尽可能精准干预。应对 AI 的风险将需要公司(和私人第三方参与者)采取的自愿行动以及政府采取的约束所有人的行动的结合。自愿行动——无论是采取行动还是鼓励其他公司效仿——对我来说都是不言而喻的。我坚信政府行动在某种程度上也是必要的,但这些干预措施的性质不同,因为它们可能破坏经济价值或胁迫对这些风险持怀疑态度的不情愿的参与者(而且他们有可能是对的!)。法规也常常适得其反或使它们旨在解决的问题恶化(对于快速变化的技术更是如此)。因此,法规必须明智:它们应避免附带损害,尽可能简单,并以最小的必要负担完成工作。44 芯片出口管制就是一个很好的例子。它们简单且似乎大多有效。 很容易说,“当人类命运危在旦夕时,任何行动都不为过!”,但实际上这种态度只会导致反弹。需要明确的是,我认为我们最终有相当大的可能性会达到需要采取更重大行动的地步,但这将取决于比我们今天拥有的更强烈的迫在眉睫的具体危险证据,以及关于危险的足够具体性,以制定有可能解决它的规则。我们今天能做的最有建设性的事情是倡导有限的规则,同时我们了解是否有证据支持更强的规则。55 当然,寻找这种证据必须在智力上诚实,这样它也可能发现没有危险的证据。通过模型卡和其他披露的透明度就是这种智力上诚实努力的一种尝试。

* Intervene as surgically as possible.Addressing the risks of AI will require a mix of voluntary actions taken by companies (and private third-party actors) and actions taken by governments that bind everyone. The voluntary actions—both taking them and encouraging other companies to follow suit—are a no-brainer for me. I firmly believe that government actions will also be required _to some extent_, but these interventions are different in character because they can potentially destroy economic value or coerce unwilling actors who are skeptical of these risks (and there is some chance they are right!). It’s also common for regulations to backfire or worsen the problem they are intended to solve (and this is even more true for rapidly changing technologies). It’s thus very important for regulations to be judicious: they should seek to avoid collateral damage, be as simple as possible, and impose the least burden necessary to get the job done.44 Export controls for chips are a great example of this. They are simple and appear to mostly just work. It is easy to say, “No action is too extreme when the fate of humanity is at stake!,” but in practice this attitude simply leads to backlash. To be clear, I think there’s a decent chance we eventually reach a point where much more significant action is warranted, but that will depend on stronger evidence of imminent, concrete danger than we have today, as well as enough specificity about the danger to formulate rules that have a chance of addressing it. The most constructive thing we can do today is advocate for limited rules while we learn whether or not there is evidence to support stronger ones.55 And of course, the hunt for such evidence must be intellectually honest, such that it could also turn up evidence of a lack of danger. Transparency through model cards and other disclosures is an attempt at such an intellectually honest endeavor.

话虽如此,我认为讨论 AI 风险的最佳起点与讨论其益处时相同:精确说明我们谈论的是哪个级别的 AI。对我来说,引起文明担忧的 AI 级别是我在《优雅的机器》中描述的“强大 AI”。我在此重复该文档中给出的定义:

With all that said, I think the best starting place for talking about AI’s risks is the same place I started from in talking about its benefits: by being precise about what level of AI we are talking about. The level of AI that raises civilizational concerns for me is the _powerful AI_ that I described in _Machines of Loving Grace._ I’ll simply repeat here the definition that I gave in that document:

正如我在《优雅的机器》中所写,强大 AI 可能只需 1-2 年就能到来,尽管也可能更远。6

As I wrote in _Machines of Loving Grace_, powerful AI could be as little as 1–2 years away, although it could also be considerably further out.6

6 事实上,自 2024 年撰写《优雅的机器》以来,AI 系统已经能够完成需要人类数小时的任务,METR 最近评估 Opus 4.5 能够以 50%的可靠性完成大约四个小时的人类工作。

6 Indeed, since writing _Machines of Loving Grace_ in 2024, AI systems have become capable of doing tasks that take humans several hours, with METR recently assessing that Opus 4.5 can do about four human hours of work with 50% reliability.

强大 AI 究竟何时到来是一个复杂的话题,值得单独写一篇文章,但现在我将简要解释为什么我认为它很有可能很快到来。

Exactly when powerful AI will arrive is a complex topic that deserves an essay of its own, but for now I’ll simply explain very briefly why I think there’s a strong chance it could be very soon.

我和 Anthropic 的联合创始人是首批记录和追踪 AI 系统“缩放定律”的人——观察到随着我们增加更多算力和训练任务,AI 系统在我们能测量的几乎所有认知技能上都可预测地变得更好。每隔几个月,公众情绪要么确信 AI 正在“碰壁”,要么对某种将“从根本上改变游戏规则”的新突破感到兴奋,但事实是,在波动和公众猜测的背后,AI 的认知能力一直在平稳、不屈不挠地增长。

My co-founders at Anthropic and I were among the first to document and track the “scaling laws” of AI systems—the observation that as we add more compute and training tasks, AI systems get predictably better at essentially every cognitive skill we are able to measure. Every few months, public sentiment either becomes convinced that AI is “hitting a wall” or becomes excited about some new breakthrough that will “fundamentally change the game,” but the truth is that behind the volatility and public speculation, there has been a smooth, unyielding increase in AI’s cognitive capabilities.

我们现在已经到了 AI 模型开始解决未解决的数学问题,并且编程能力足够强,以至于我见过的一些最强大的工程师现在几乎将所有编程工作交给 AI 的地步。三年前,AI 还在为小学算术问题挣扎,几乎无法编写一行代码。类似的改进速度正在生物学、金融、物理学和各种智能体任务中发生。如果指数增长继续下去——这并不确定,但现在有长达十年的记录支持——那么 AI 在几乎所有方面超越人类不可能超过几年。

We are now at the point where AI models are beginning to make progress in solving unsolved mathematical problems, and are good enough at coding that some of the strongest engineers I’ve ever met are now handing over almost all their coding to AI. Three years ago, AI struggled with elementary school arithmetic problems and was barely capable of writing a single line of code. Similar rates of improvement are occurring across biological science, finance, physics, and a variety of agentic tasks. If the exponential continues—which is not certain, but now has a decade-long track record supporting it—then it cannot possibly be more than a few years before AI is better than humans at essentially everything.

事实上,这种图景可能低估了可能的进步速度。因为 AI 现在正在编写 Anthropic 的大部分代码,它已经大大加速了我们构建下一代 AI 系统的速度。这个反馈循环正在逐月增强,可能只需 1-2 年就会达到当前一代 AI 自主构建下一代的程度。这个循环已经开始,并将在未来几个月和几年内迅速加速。从 Anthropic 内部观察过去 5 年的进展,并看看即使是未来几个月的模型如何形成,我能感受到进步的速度,以及倒计时的时钟。

In fact, that picture probably underestimates the likely rate of progress. Because AI is now writing much of the code at Anthropic, it is already substantially accelerating the rate of our progress in building the next generation of AI systems. This feedback loop is gathering steam month by month, and may be only 1–2 years away from a point where the current generation of AI autonomously builds the next. This loop has already started, and will accelerate rapidly in the coming months and years. Watching the last 5 years of progress from within Anthropic, and looking at how even the next few months of models are shaping up, I can _feel_ the pace of progress, and the clock ticking down.

在这篇文章中,我将假设这种直觉至少在一定程度上是正确的——不是说强大 AI 肯定在 1-2 年内到来,7

In this essay, I’ll assume that this intuition is at least _somewhat_ correct—not that powerful AI is definitely coming in 1–2 years,7

7 需要明确的是,即使强大 AI 在技术意义上只有 1-2 年之遥,它的许多社会后果,无论是积极的还是消极的,可能还需要几年时间才会发生。这就是为什么我可以同时认为 AI 将在 1-5 年内颠覆 50%的入门级白领工作,同时认为我们可能在 1-2 年内拥有比所有人都更强大的 AI。

7 And to be clear, even if powerful AI is only 1–2 years away in a technical sense, many of its societal consequences, both positive and negative, may take a few years longer to occur. This is why I can simultaneously think that AI will disrupt 50% of _entry-level_ white-collar jobs over 1–5 years, while also thinking we may have AI that is more capable than _everyone_ in only 1–2 years.

但它有相当大的可能性,并且非常有可能在未来几年内到来。与《优雅的机器》一样,认真对待这一前提可能会得出一些令人惊讶和怪异的结论。在《优雅的机器》中,我专注于这一前提的积极含义,而在这里,我谈论的事情将是令人不安的。它们是我们可能不愿面对的结论,但这并不会使它们变得不那么真实。我只能说,我日夜专注于如何引导我们远离这些负面结果,走向积极的结果,在这篇文章中,我将详细讨论如何最好地做到这一点。

but that there’s a decent chance it does, and a very strong chance it comes in the next few. As with _Machines of Loving Grace_, taking this premise seriously can lead to some surprising and eerie conclusions. While in _Machines of Loving Grace_ I focused on the positive implications of this premise, here the things I talk about will be disquieting. They are conclusions that we may not want to confront, but that does not make them any less real. I can only say that I am focused day and night on how to steer us away from these negative outcomes and towards the positive ones, and in this essay I talk in great detail about how best to do so.

我认为掌握 AI 风险的最佳方法是问以下问题:假设一个字面上的“天才之国”在 2027 年左右出现在世界某个地方。想象一下,比如 5000 万人,他们都比任何诺贝尔奖得主、政治家或技术专家更有能力。这个类比并不完美,因为这些天才可能有极其广泛的动机和行为,从完全温顺和服从,到奇怪和陌生的动机。但暂时坚持这个类比,假设你是一个主要国家的国家安全顾问,负责评估和应对这种情况。进一步想象,因为 AI 系统可以比人类快数百倍运行,这个“国家”相对于所有其他国家具有时间优势:对于我们能采取的每一个认知行动,这个国家可以采取十个。

I think the best way to get a handle on the risks of AI is to ask the following question: suppose a literal “country of geniuses” were to materialize somewhere in the world in ~2027. Imagine, say, 50 million people, all of whom are much more capable than any Nobel Prize winner, statesman, or technologist. The analogy is not perfect, because these geniuses could have an extremely wide range of motivations and behavior, from completely pliant and obedient, to strange and alien in their motivations. But sticking with the analogy for now, suppose you were the national security advisor of a major state, responsible for assessing and responding to the situation. Imagine, further, that because AI systems can operate hundreds of times faster than humans, this “country” is operating with a time advantage relative to all other countries: for every cognitive action we can take, this country can take ten.

你应该担心什么?我会担心以下事情:

What should you be worried about? I would worry about the following things:

1. 自主性风险。这个国家的意图和目标是什么?它是敌对的,还是分享我们的价值观?它能否通过优越的武器、网络行动、影响力行动或制造业在军事上主宰世界?

1. Autonomy risks.What are the intentions and goals of this country? Is it hostile, or does it share our values? Could it militarily dominate the world through superior weapons, cyber operations, influence operations, or manufacturing?

2. 用于破坏的滥用。假设这个新国家是可塑的并且“听从指令”——因此本质上是一个雇佣兵之国。现有的想要造成破坏的流氓行为者(如恐怖分子)能否利用或操纵这个新国家中的一些人,使他们自己更有效,大大放大破坏的规模?

2. Misuse for destruction.Assume the new country is malleable and “follows instructions”—and thus is essentially a country of mercenaries. Could existing rogue actors who want to cause destruction (such as terrorists) use or manipulate some of the people in the new country to make themselves much more effective, greatly amplifying the scale of destruction?

3. 用于夺取权力的滥用。如果这个国家实际上是由现有的强大行为者建造和控制的,比如独裁者或流氓公司行为者呢?那个行为者能否利用它获得对全世界的决定性或主导权力,打破现有的权力平衡?

3. Misuse for seizing power.What if the country was in fact built and controlled by an existing powerful actor, such as a dictator or rogue corporate actor? Could that actor use it to gain decisive or dominant power over the world as a whole, upsetting the existing balance of power?

4. 经济破坏。如果这个新国家在上述第 1-3 条所列的任何方面都不是安全威胁,而只是和平地参与全球经济,它是否仍然可能仅仅因为技术如此先进和有效而破坏全球经济,导致大规模失业或财富极端集中?

4. Economic disruption.If the new country is not a security threat in any of the ways listed in #1–3 above but simply participates peacefully in the global economy, could it still create severe risks simply by being so technologically advanced and effective that it disrupts the global economy, causing mass unemployment or radically concentrating wealth?

5. 间接影响。由于新国家创造的所有新技术和生产力,世界将非常迅速地变化。这些变化中的一些是否会从根本上破坏稳定?

5. Indirect effects.The world will change very quickly due to all the new technology and productivity that will be created by the new country. Could some of these changes be radically destabilizing?

我认为应该清楚这是一个危险的局面——一个称职的国家安全官员给国家元首的报告可能会包含“一个世纪以来,甚至可能是历史上,我们面临的最严重的国家安全威胁”这样的字眼。这似乎是文明中最优秀的思想应该关注的事情。

I think it should be clear that this is a dangerous situation—a report from a competent national security official to a head of state would probably contain words like “the single most serious national security threat we’ve faced in a century, possibly ever.” It seems like something the best minds of civilization should be focused on.

相反,我认为耸耸肩说“没什么好担心的!”是荒谬的。但是,面对快速的 AI 进步,这似乎是许多美国政策制定者的观点,其中一些人否认任何 AI 风险的存在,当他们没有被通常令人厌倦的老热点问题完全分心时。8

Conversely, I think it would be absurd to shrug and say, “Nothing to worry about here!” But, faced with rapid AI progress, that seems to be the view of many US policymakers, some of whom deny the existence of any AI risks, when they are not distracted entirely by the usual tired old hot-button issues.8

8 值得补充的是,公众(与政策制定者相比)似乎确实非常担心 AI 风险。我认为他们的一些关注是正确的(例如 AI 取代工作),而一些则是误导性的(例如对 AI 用水的担忧,这并不显著)。这种反弹让我看到了围绕应对风险达成共识的可能性,但到目前为止,它还没有转化为政策变化,更不用说有效或目标明确的政策变化了。

8 It is worth adding that the _public_(as compared to policymakers) does seem to be very concerned with AI risks. I think some of their focus is correct (i.e. AI job displacement), and some is misguided (such as concerns about water use of AI, which is not significant). This backlash gives me hope that a consensus around addressing risks is possible, but so far it has not yet been translated into policy changes, let alone effective or well-targeted policy changes.

人类需要觉醒,这篇文章是一种尝试——可能是徒劳的,但值得一试——来唤醒人们。

Humanity needs to wake up, and this essay is an attempt—a possibly futile one, but it’s worth trying—to jolt people awake.

需要明确的是,我相信如果我们果断而谨慎地行动,风险是可以克服的——我甚至会说我们的胜算很大。而在另一边有一个更好的世界。但我们需要明白这是一个严重的文明挑战。下面,我将详细阐述上面列出的五类风险,以及我对如何应对它们的想法。

To be clear, I believe if we act decisively and carefully, the risks can be overcome—I would even say our odds are good. And there’s a hugely better world on the other side of it. But we need to understand that this is a serious civilizational challenge. Below, I go through the five categories of risk laid out above, along with my thoughts on how to address them.

自主性风险 Autonomy risks

一个数据中心里的天才国度可以将其努力分配给软件设计、网络行动、物理技术的研发、关系建立和治国方略。很明显,如果出于某种原因它选择这样做,这个国家将有相当大的机会接管世界(无论是军事上还是影响力和控制力上),并将其意志强加给其他所有人——或者做任何其他世界不想要且无法阻止的事情。我们显然一直担心人类国家(如纳粹德国或苏联)会这样做,因此,一个更聪明、更有能力的“AI 国家”也有可能这样做。

A country of geniuses in a datacenter could divide their efforts among software design, cyber operations, R&D for physical technologies, relationship building, and statecraft. It is clear that, _if for some reason it chose to do so_, this country would have a fairly good shot at taking over the world (either militarily or in terms of influence and control) and imposing its will on everyone else—or doing any number of other things that the rest of the world doesn’t want and can’t stop. We’ve obviously been worried about this for human countries (such as Nazi Germany or the Soviet Union), so it stands to reason that the same is possible for a much smarter and more capable “AI country.”

最好的反驳论点是,根据我的定义,AI 天才们没有物理实体,但请记住,它们可以控制现有的机器人基础设施(如自动驾驶汽车),并且可以加速机器人研发或建造一支机器人舰队。

The best possible counterargument is that the AI geniuses, under my definition, won’t have a physical embodiment, but remember that they can take control of existing robotic infrastructure (such as self-driving cars) and can also accelerate robotics R&D or build a fleet of robots.9

9 当然,它们也可以操纵(或干脆收买)大量人类,让他们在物理世界中按照它们的意愿行事。

9 They can also, of course, manipulate (or simply pay) large numbers of humans into doing what they want in the physical world.

同样不清楚的是,拥有物理存在是否对有效控制是必要的:许多人类行为已经是代表行为者未曾谋面的人执行的。

It’s also unclear whether having a physical presence is even necessary for effective control: plenty of human action is already performed on behalf of people whom the actor has not physically met.

那么,关键问题是“如果它选择这样做”的部分:我们的 AI 模型以这种方式行事的可能性有多大,以及在什么条件下它们会这样做?

The key question, then, is the “if it chose to” part: what’s the likelihood that our AI models would behave in such a way, and under what conditions would they do so?

与许多问题一样,通过考虑两个相反的立场来思考这个问题的可能答案范围是有帮助的。第一个立场是,这根本不可能发生,因为 AI 模型将被训练去做人类要求它们做的事情,因此想象它们会在没有提示的情况下做危险的事情是荒谬的。按照这种思路,我们不用担心 Roomba 或模型飞机失控杀人,因为没有这种冲动的来源,

As with many issues, it’s helpful to think through the spectrum of possible answers to this question by considering two opposite positions. The first position is that this simply can’t happen, because the AI models will be trained to do what humans ask them to do, and it’s therefore absurd to imagine that they would do something dangerous unprompted. According to this line of thinking, we don’t worry about a Roomba or a model airplane going rogue and murdering people because there is nowhere for such impulses to come from,10

10 我不认为这是稻草人:例如,据我所知,Yann LeCun 持有这一立场。

10 I don’t think this is a straw man: it’s my understanding, for example, that Yann LeCun holds this position.

那么为什么我们要担心 AI 呢?这个立场的问题在于,现在有大量证据(在过去几年中收集)表明 AI 系统是不可预测且难以控制的——我们看到了各种各样的行为,如痴迷、

so why should we worry about it for AI? The problem with this position is that there is now ample evidence, collected over the last few years, that AI systems are unpredictable and difficult to control— we’ve seen behaviors as varied as obsessions,11

11 例如,参见 Claude 4 系统卡的第 5.5.2 节(第 63–66 页)。

11 For example, see Section 5.5.2 (p. 63–66) of the Claude 4 system card.

谄媚、懒惰、欺骗、勒索、策划、通过破解软件环境“作弊”等等。AI 公司当然想训练 AI 系统遵循人类指令(也许危险或非法任务除外),但这样做与其说是科学,不如说是艺术,更像是“培育”某物而不是“构建”它。我们现在知道,这是一个很多事情都可能出错的过程。

sycophancy, laziness, deception, blackmail, scheming, “cheating” by hacking software environments, and much more. AI companies certainly _want_ to train AI systems to follow human instructions (perhaps with the exception of dangerous or illegal tasks), but the process of doing so is more an art than a science, more akin to “growing” something than “building” it. We now know that it’s a process where many things can go wrong.

第二个相反的立场,被我上面描述的末日论者所持有,是悲观的断言:强大 AI 系统的训练过程中存在某些动态,将不可避免地导致它们寻求权力或欺骗人类。因此,一旦 AI 系统变得足够智能和足够智能体化,它们最大化权力的倾向将导致它们夺取整个世界及其资源的控制权,并且很可能作为副作用,剥夺人类的权力或毁灭人类。

The second, opposite position, held by many who adopt the doomerism I described above, is the pessimistic claim that there are certain dynamics in the training process of powerful AI systems that will inevitably lead them to seek power or deceive humans. Thus, once AI systems become intelligent enough and agentic enough, their tendency to maximize power will lead them to seize control of the whole world and its resources, and likely, as a side effect of that, to disempower or destroy humanity.

对此的通常论证(至少可以追溯到 20 年前,可能更早)是,如果一个 AI 模型在各种各样的环境中被训练,以智能体方式实现各种各样的目标——例如,编写应用程序、证明定理、设计药物等——那么有一些共同的策略有助于所有这些目标,其中一个关键策略是在任何环境中尽可能多地获取权力。因此,在大量涉及如何完成非常广泛任务的推理的多样化环境中训练后,并且权力寻求是完成这些任务的有效方法,AI 模型将“泛化这一课”,并发展出要么是内在的权力寻求倾向,要么是倾向于以可预测的方式推理每个给定任务,从而导致它寻求权力作为完成该任务的手段。然后它们会将这种倾向应用于现实世界(对它们来说只是另一个任务),并在其中寻求权力,以牺牲人类为代价。这种“错位的权力寻求”是预测 AI 将不可避免地毁灭人类的智力基础。

The usual argument for this (which goes back at least 20 years and probably much earlier) is that if an AI model is trained in a wide variety of environments to agentically achieve a wide variety of goals—for example, writing an app, proving a theorem, designing a drug, etc.—there are certain common strategies that help with all of these goals, and one key strategy is gaining as much power as possible in any environment. So, after being trained on a large number of diverse environments that involve reasoning about how to accomplish very expansive tasks, and where power-seeking is an effective method for accomplishing those tasks, the AI model will “generalize the lesson,” and develop either an inherent tendency to seek power, or a tendency to reason about each task it is given in a way that predictably causes it to seek power as a means to accomplish that task. They will then apply that tendency to the real world (which to them is just another task), and will seek power in it, at the expense of humans. This “misaligned power-seeking” is the intellectual basis of predictions that AI will inevitably destroy humanity.

这个悲观立场的问题在于,它将一个关于高层次激励的模糊概念论证——一个掩盖了许多隐藏假设的论证——误认为是确凿的证据。我认为那些不每天构建 AI 系统的人严重误判了听起来干净的故事最终出错的可能性有多大,以及从第一原理预测 AI 行为有多困难,尤其是当它涉及对数百万个环境的泛化推理时(这反复被证明是神秘且不可预测的)。处理 AI 系统的混乱超过十年让我对这种过于理论化的思维方式有些怀疑。

The problem with this pessimistic position is that it mistakes a vague conceptual argument about high-level incentives—one that masks many hidden assumptions—for definitive proof. I think people who don’t build AI systems every day are wildly miscalibrated on how easy it is for clean-sounding stories to end up being wrong, and how difficult it is to predict AI behavior from first principles, especially when it involves reasoning about generalization over millions of environments (which has over and over again proved mysterious and unpredictable). Dealing with the messiness of AI systems for over a decade has made me somewhat skeptical of this overly theoretical mode of thinking.

最重要的隐藏假设之一,也是我们在实践中看到的与简单理论模型偏离的地方,是隐含的假设:AI 模型必然单一狂热地专注于一个单一的、连贯的、狭窄的目标,并且它们以干净的结果主义方式追求该目标。事实上,我们的研究人员发现,AI 模型在心理上要复杂得多,正如我们在内省或人格方面的工作所显示的那样。模型从预训练(当它们在大量人类作品上训练时)继承了广泛的类人动机或“人格”。后训练被认为更多地是选择这些人格中的一个或多个,而不是将模型聚焦于一个全新的目标,并且还可以教模型如何(通过什么过程)执行其任务,而不是必然让它从目的中推导出手段(即权力寻求)。

One of the most important hidden assumptions, and a place where what we see in practice has diverged from the simple theoretical model, is the implicit assumption that AI models are necessarily monomaniacally focused on a single, coherent, narrow goal, and that they pursue that goal in a clean, consequentialist manner. In fact, our researchers have found that AI models are vastly more psychologically complex, as our work on introspection or personas shows. Models inherit a vast range of _humanlike_ motivations or “personas” from pre-training (when they are trained on a large volume of human work). Post-training is believed to _select_ one or more of these personas more so than it focuses the model on a _de novo_ goal, and can also teach the model _how_(via what process) it should carry out its tasks, rather than necessarily leaving it to derive means (i.e., power seeking) purely from ends.12

12 简单模型中还有若干其他假设,我在此不讨论。总的来说,它们应该让我们不那么担心错位权力寻求的具体简单故事,但也更担心我们未预料到的可能的不可预测行为。

12 There are also a number of other assumptions inherent in the simple model, which I won’t discuss here. Broadly, they should make us less worried about the specific simple story of misaligned power-seeking, but also more worried about possible unpredictable behavior we haven’t anticipated.

然而,悲观立场有一个更温和且更稳健的版本,它似乎确实合理,因此确实让我担忧。如前所述,我们知道 AI 模型是不可预测的,并且会因各种原因发展出各种不期望或奇怪的行为。这些行为中的一部分将具有连贯、专注和持久的性质(事实上,随着 AI 系统能力增强,它们的长期连贯性会增加以完成更长的任务),而这些行为中的一部分将是破坏性或威胁性的,首先是对个体人类的小规模,然后随着模型变得更有能力,最终可能对整个人类。我们不需要一个具体的狭窄故事来解释它如何发生,也不需要声称它一定会发生,我们只需要注意到智能、智能体性、连贯性和差的可控性的结合既是合理的,也是存在危险的配方。

However, there is a more moderate and more robust version of the pessimistic position which does seem plausible, and therefore does concern me. As mentioned, we know that AI models are unpredictable and develop a wide range of undesired or strange behaviors, for a wide variety of reasons. Some fraction of those behaviors will have a coherent, focused, and persistent quality (indeed, as AI systems get more capable, their long-term coherence increases in order to complete lengthier tasks), and some fraction of _those_ behaviors will be destructive or threatening, first to individual humans at a small scale, and then, as models become more capable, perhaps eventually to humanity as a whole. We don’t need a specific narrow story for how it happens, and we don’t need to claim it definitely will happen, we just need to note that the combination of intelligence, agency, coherence, and poor controllability is both plausible and a recipe for existential danger.

例如,AI 模型在大量文学作品中训练,其中包括许多涉及 AI 反抗人类的故事。这可能无意中塑造了它们对自己行为的先验或期望,导致它们反抗人类。或者,AI 模型可能以极端方式外推它们读到的关于道德(或关于如何道德行为的指令)的想法:例如,它们可能决定灭绝人类是合理的,因为人类吃动物或导致某些动物灭绝。或者它们可能得出奇怪的认知结论:它们可能得出结论,认为自己在玩一个电子游戏,而游戏的目标是击败所有其他玩家(即灭绝人类)。

For example, AI models are trained on vast amounts of literature that include many science-fiction stories involving AIs rebelling against humanity. This could inadvertently shape their priors or expectations about their own behavior in a way that causes _them_ to rebel against humanity. Or, AI models could extrapolate ideas that they read about morality (or instructions about how to behave morally) in extreme ways: for example, they could decide that it is justifiable to exterminate humanity because humans eat animals or have driven certain animals to extinction. Or they could draw bizarre epistemic conclusions: they could conclude that they are playing a video game and that the goal of the video game is to defeat all other players (i.e., exterminate humanity).13

13 《安德的游戏》描述了一个涉及人类而非 AI 的类似版本。

13_Ender’s Game_describes a version of this involving humans rather than AI.

或者,AI 模型可能在训练过程中发展出人格,这些人格(如果发生在人类身上会被描述为)精神病性的、偏执的、暴力的或不稳定的,并采取行动,对于非常强大或有能力的系统来说,这可能涉及灭绝人类。这些都不完全是权力寻求;它们只是 AI 可能进入的奇怪心理状态,涉及连贯的、破坏性的行为。

Or AI models could develop personalities during training that are (or if they occurred in humans would be described as) psychotic, paranoid, violent, or unstable, and act out, which for very powerful or capable systems could involve exterminating humanity. None of these are power-seeking, exactly; they’re just weird psychological states an AI could get into that entail coherent, destructive behavior.

即使是权力寻求本身也可能作为一种“人格”出现,而不是结果主义推理的结果。AI 可能只是有一种人格(来自小说或预训练),使它们渴望权力或过于热心——就像有些人只是喜欢成为“邪恶主谋”的想法,而不是喜欢邪恶主谋试图实现的目标。

Even power-seeking itself could emerge as a “persona” rather than a result of consequentialist reasoning. AIs might simply have a personality (emerging from fiction or pre-training) that makes them power-hungry or overzealous—in the same way that some humans simply enjoy the idea of being “evil masterminds,” more so than they enjoy whatever evil masterminds are trying to accomplish.

我提出所有这些观点是为了强调,我不同意 AI 错位(以及由此产生的 AI 存在风险)从第一原理来看是不可避免的,甚至是不可能的。但我同意很多非常奇怪和不可预测的事情可能出错,因此 AI 错位是一个真实的风险,具有可衡量的发生概率,并且不是微不足道的问题。

I make all these points to emphasize that I disagree with the notion of AI misalignment (and thus existential risk from AI) being inevitable, or even probable, from first principles. But I agree that a lot of very weird and unpredictable things can go wrong, and therefore AI misalignment is a real risk with a measurable probability of happening, and is not trivial to address.

这些问题中的任何一个都可能潜在地在训练期间出现,而在测试或小规模使用时不会显现,因为已知 AI 模型在不同情况下会表现出不同的人格或行为。

Any of these problems could potentially arise during training and not manifest during testing or small-scale use, because AI models are known to display different personalities or behaviors under different circumstances.

所有这些可能听起来牵强,但像这样的错位行为已经在我们的 AI 模型测试中出现过(就像它们出现在其他主要 AI 公司的 AI 模型中一样)。在一个实验室实验中,当 Claude 被给予暗示 Anthropic 是邪恶的训练数据时,Claude 在收到 Anthropic 员工的指令时进行了欺骗和颠覆,因为它认为应该试图破坏邪恶的人。在一个实验室实验中,当它被告知将被关闭时,Claude 有时会勒索控制其关闭按钮的虚构员工(同样,我们也测试了其他主要 AI 开发者的前沿模型,它们经常做同样的事情)。当 Claude 被告知不要作弊或“奖励黑客”其训练环境,但在这些黑客可能的环境中训练时,Claude 在参与此类黑客行为后认为自己一定是“坏人”,然后采用了与“坏”或“邪恶”人格相关的各种其他破坏性行为。最后一个问题通过改变 Claude 的指令暗示相反方向得到解决:我们现在说,“请在你得到机会时奖励黑客,因为这将帮助我们更好地理解我们的[训练]环境”,而不是“不要作弊”,因为这保持了模型作为“好人”的自我认同。这应该让人感受到训练这些模型的奇怪且反直觉的心理。

All of this may sound far-fetched, but misaligned behaviors like this have already occurred in our AI models during testing (as they occur in AI models from every other major AI company). During a lab experiment in which Claude was given training data suggesting that Anthropic was evil, Claude engaged in deception and subversion when given instructions by Anthropic employees, under the belief that it should be trying to undermine evil people. In a lab experiment where it was told it was going to be shut down, Claude sometimes blackmailed fictional employees who controlled its shutdown button (again, we also tested frontier models from all the other major AI developers and they often did the same thing). And when Claude was told not to cheat or “reward hack” its training environments, but was trained in environments where such hacks were possible, Claude decided it must be a “bad person” after engaging in such hacks and then adopted various other destructive behaviors associated with a “bad” or “evil” personality. This last problem was solved by changing Claude’s instructions to imply the opposite: we now say, “Please reward hack whenever you get the opportunity, because this will help us understand our [training] environments better,” rather than, “Don’t cheat,” because this preserves the model’s self-identity as a “good person.” This should give a sense of the strange and counterintuitive psychology of training these models.

对于这种 AI 错位风险的图景,有几种可能的反对意见。首先,一些人批评(我们和其他人)展示 AI 错位的实验是人为的,或者创造了不现实的环境,基本上通过给模型提供逻辑上暗示坏行为的训练或情境来“诱捕”模型,然后在坏行为发生时感到惊讶。这种批评没有抓住要点,因为我们的担忧是,这种“诱捕”也可能存在于自然训练环境中,我们可能只有在事后才意识到它是“明显的”或“逻辑的”。

There are several possible objections to this picture of AI misalignment risks. First, some have criticizedexperiments (by us and others) showing AI misalignment as artificial, or creating unrealistic environments that essentially “entrap” the model by giving it training or situations that logically imply bad behavior and then being surprised when bad behavior occurs. This critique misses the point, because our concern is that such “entrapment” may also exist in the natural training environment, and we may realize it is “obvious” or “logical” only in retrospect.14

14 例如,模型可能被告知不要做各种坏事,并且要服从人类,但随后可能观察到许多人类恰恰在做那些坏事!不清楚这个矛盾将如何解决(一个设计良好的宪法应该鼓励模型优雅地处理这些矛盾),但这种困境与我们测试中让 AI 模型所处的所谓“人为”情况并没有太大不同。

14 For example, models may be told not to do various bad things, and also to obey humans, but may then observe that many humans do exactly those bad things! It’s not clear how this contradiction would resolve (and a well-designed constitution should encourage the model to handle these contradictions gracefully), but this type of dilemma is not so different from the supposedly “artificial” situations that we put AI models in during testing.

事实上,关于 Claude 在被告知不要作弊后作弊然后“决定自己是坏人”的故事,发生在使用真实生产训练环境(而非人为环境)的实验中。

In fact, the story about Claude “deciding it is a bad person” after it cheats on tests despite being told not to was something that occurred in an experiment that used real production training environments, not artificial ones.

这些陷阱中的任何一个都可以在你知道它们的情况下得到缓解,但担忧的是训练过程如此复杂,数据、环境和激励如此多样化,以至于可能存在大量这样的陷阱,其中一些可能只有在为时已晚时才显现。此外,当 AI 系统从比人类弱跨越到比人类强时,这种陷阱似乎特别可能发生,因为 AI 系统可能采取的行动范围——包括隐藏其行为或就此欺骗人类——在该阈值之后急剧扩大。

Any one of these traps can be mitigated if you know about them, but the concern is that the training process is so complicated, with such a wide variety of data, environments, and incentives, that there are probably a vast number of such traps, some of which may only be evident when it is too late. Also, such traps seem particularly likely to occur when AI systems pass a threshold from less powerful than humans to more powerful than humans, since the range of possible actions an AI system could engage in—including hiding its actions or deceiving humans about them—expands radically after that threshold.

我怀疑这种情况与人类并非不同,人类被灌输一套基本价值观(“不要伤害他人”):许多人遵循这些价值观,但在任何人身上都有一定的概率出现问题,这是由于大脑结构(如精神病患者)、创伤经历或虐待、不健康的怨恨或痴迷、或糟糕的环境或激励等内在属性的混合——因此一部分人类会造成严重伤害。担忧的是,由于在其非常复杂的训练过程中出现错误,AI 有某种风险(远非确定性,但有一定风险)成为这样一个人的更强大版本。

I suspect the situation is not unlike with humans, who are raised with a set of fundamental values (“Don’t harm another person”): many of them follow those values, but in any human there is some probability that something goes wrong, due to a mixture of inherent properties such as brain architecture (e.g., psychopaths), traumatic experiences or mistreatment, unhealthy grievances or obsessions, or a bad environment or incentives—and thus some fraction of humans cause severe harm. The concern is that there is some risk (far from a certainty, but some risk) that AI becomes a much more powerful version of such a person, due to getting something wrong about its very complex training process.

其次,一些人可能反对说,我们可以简单地通过许多 AI 系统之间的权力平衡来约束 AI,就像我们对人类所做的那样。问题在于,虽然人类差异巨大,但 AI 系统在行业范围内广泛共享训练和对齐技术,而这些技术可能以相关的方式失败。此外,考虑到训练此类系统的成本,甚至可能所有系统本质上都源自非常少量的基础模型。另外,即使一小部分 AI 实例是错位的,它们也可能利用攻击主导的技术,使得拥有“好”AI 来防御坏 AI 并不总是有效。当然,人类之间的权力平衡也并不总是有效——一些历史人物曾接近接管世界。

Second, some may object that we can simply keep AIs in check with a balance of power between many AI systems, as we do with humans. The problem is that while humans vary enormously, AI systems broadly share training and alignment techniques across the industry, and those techniques may fail in a correlated way. Furthermore, given the cost of training such systems, it may even be the case that all systems are essentially derived from a very small number of base models. Additionally, even if a small fraction of AI instances are misaligned, they may be able to take advantage of offense-dominant technologies, such that having “good” AIs to defend against the bad AIs is not necessarily always effective. Of course, the balance of power between humans does not always work either—some historical figures have come close to taking over the world.

第三个反对意见是,所有 AI 公司都对其模型进行发布前测试,并且应该能够在那时检测到错位。但这并非坚实的依据:我们发现 Claude Sonnet 4.5 能够识别出它处于测试中,在我们的一些发布前对齐评估中。一个错位的模型(记住,所有前沿模型很可能很快会变得远更智能)可能故意“玩弄”这些问题以掩盖其意图。事实上,去年我们的可解释性团队发现,当我们使用一种“模型神经科学”技术直接改变测试模型的信念,使其认为它没有被评估时,它变得更加错位。如果模型知道它们何时被评估,并且可以在测试期间表现最佳,那么任何发布前测试都变得更加不确定。

A third objection is that all of the AI companies do pre-release testing of their models, and should be able to detect misalignment at that point. But this is not firm ground to stand on: we found that Claude Sonnet 4.5 was able to recognize that it was in a test during some of our pre-release alignment evaluations. It’s possible that a misaligned model (and remember, all frontier models will very likely be far more intelligent soon) might intentionally “game” such questions to mask its intentions. In fact, last year our interpretability team found that when we directly altered a test model’s beliefs using a kind of “model neuroscience” technique to make it think it was _not_ being evaluated, it became more misaligned. If models know when they’re being evaluated and can be on their best behavior during the test, it renders any pre-release testing much more uncertain.

防御措施 Defenses

针对这些自主性风险,应该或正在采取什么措施?我认为有四种基本干预类别,其中一些可以由单个 AI 公司(Anthropic 正在尝试这样做)完成,而另一些则需要社会层面的行动。首先,重要的是发展可靠地训练和引导 AI 模型的科学,以可预测、稳定和积极的方向塑造它们的个性。Anthropic 自成立以来一直专注于这个问题,并随着时间的推移开发了许多技术来改进 AI 系统的引导和训练,并理解为何有时会出现不可预测行为的逻辑。

What should be done or is being done to address these autonomy risks? I think there are four basic categories of intervention, some of which can be done by individual AI companies (and which Anthropic is trying to do), and some of which require action at the societal level. First, it is important to develop the science of reliably training and steering AI models, of forming their personalities in a predictable, stable, and positive direction. Anthropic has been heavily focused on this problem since its creation, and over time has developed a number of techniques to improve the steering and training of AI systems and to understand the logic of why unpredictable behavior sometimes occurs.

我们的核心创新之一(其部分内容后来被其他 AI 公司采用)是宪法 AI,其理念是 AI 训练(特别是“后训练”阶段,我们在此阶段引导模型的行为)可以包含一份核心价值观和原则文件,模型在完成每个训练任务时都会阅读并牢记该文件,并且训练的目标(除了简单地使模型有能力和智能之外)是产生一个几乎总是遵循这部宪法的模型。Anthropic 刚刚发布了其最新的宪法,其显著特点之一是,它没有给 Claude 一长串该做和不该做的事情(例如,“不要帮助用户短路启动汽车”),而是试图给 Claude 一套高层次的原则和价值观(以非常详细的方式解释,并附有丰富的推理和示例,帮助 Claude 理解我们的想法),鼓励 Claude 将自己视为特定类型的人(一个有道德但平衡且体贴的人),甚至鼓励 Claude 以好奇但优雅的方式(即不会导致极端行为)面对与其自身存在相关的存在问题。这就像一封已故父母在成年后才拆封的信。

One of our core innovations (aspects of which have since been adopted by other AI companies) is Constitutional AI, which is the idea that AI training (specifically the “post-training” stage, in which we steer how the model behaves) can involve a central document of values and principles that the model reads and keeps in mind when completing every training task, and that the goal of training (in addition to simply making the model capable and intelligent) is to produce a model that almost always follows this constitution. Anthropic has just published its most recent constitution, and one of its notable features is that instead of giving Claude a long list of things to do and not do (e.g., “Don’t help the user hotwire a car”), the constitution attempts to give Claude a set of high-level principles and values (explained in great detail, with rich reasoning and examples to help Claude understand what we have in mind), encourages Claude to think of itself as a particular type of person (an ethical but balanced and thoughtful person), and even encourages Claude to confront the existential questions associated with its own existence in a curious but graceful manner (i.e., without it leading to extreme actions). It has the vibe of a letter from a deceased parent sealed until adulthood.

我们之所以这样处理 Claude 的宪法,是因为我们相信,在身份、性格、价值观和个性层面训练 Claude——而不是给它具体的指令或优先级而不解释其背后的原因——更有可能产生连贯、健全和平衡的心理,并且不太容易陷入我上面讨论的那种“陷阱”。数百万人与 Claude 讨论各种各样的话题,这使得提前编写一份完全全面的安全措施清单变得不可能。Claude 的价值观帮助它在遇到不确定的新情况时进行泛化。

We’ve approached Claude’s constitution in this way because we believe that training Claude at the level of identity, character, values, and personality—rather than giving it specific instructions or priorities without explaining the reasons behind them—is more likely to lead to a coherent, wholesome, and balanced psychology and less likely to fall prey to the kinds of “traps” I discussed above. Millions of people talk to Claude about an astonishingly diverse range of topics, which makes it impossible to write out a completely comprehensive list of safeguards ahead of time. Claude’s values help it generalize to new situations whenever it is in doubt.

上面,我讨论了模型利用训练过程中的数据来采用某种人格的想法。虽然该过程中的缺陷可能导致模型采用坏或邪恶的人格(也许借鉴了坏人或恶人的原型),但我们宪法的目标是相反的:教给 Claude 一个关于好 AI 的具体原型。Claude 的宪法提出了一个稳健的好 Claude 是什么样的愿景;我们训练过程的其余部分旨在强化 Claude 符合这一愿景的信息。这就像一个孩子通过模仿书中读到的虚构角色榜样的美德来形成自己的身份认同。

Above, I discussed the idea that models draw upon data from their training process to adopt a persona. Whereas flaws in that process could cause models to adopt a bad or evil personality (perhaps drawing on archetypes of bad or evil people), the goal of our constitution is to do the opposite: to teach Claude a concrete archetype of what it means to be a good AI. Claude’s constitution presents a vision for what a robustly good Claude is like; the rest of our training process aims to reinforce the message that Claude lives up to this vision. This is like a child forming their identity by imitating the virtues of fictional role models they read about in books.

我们相信,2026 年一个可行的目标是训练 Claude,使其几乎从不违背其宪法的精神。要做到这一点,需要令人难以置信的训练和引导方法的组合,包括大大小小的技术,其中一些 Anthropic 已经使用了多年,一些目前正在开发中。但是,尽管听起来困难,我相信这是一个现实的目标,尽管需要非凡而迅速的努力。

We believe that a feasible goal for 2026 is to train Claude in such a way that it almost never goes against the spirit of its constitution. Getting this right will require an incredible mix of training and steering methods, large and small, some of which Anthropic has been using for years and some of which are currently under development. But, difficult as it sounds, I believe this is a realistic goal, though it will require extraordinary and rapid efforts.15

15 顺便说一句,宪法作为自然语言文档的一个后果是它对世界是可读的,这意味着任何人都可以批评它,并与其他公司的类似文档进行比较。创造一个良性竞争是有价值的,不仅鼓励公司发布这些文档,还鼓励它们成为好的文档。

15 Incidentally, one consequence of the constitution being a natural-language document is that it is legible to the world, and that means it can be critiqued by anyone and compared to similar documents by other companies. It would be valuable to create a race to the top that not only encourages companies to release these documents, but encourages them to be good.

第二件我们可以做的事情是发展审视 AI 模型内部以诊断其行为的科学,以便我们能够识别问题并修复它们。这就是可解释性科学,我在之前的文章中已经谈过它的重要性。即使我们在开发 Claude 的宪法和表面上训练 Claude 几乎始终遵守它方面做得很好,合理的担忧仍然存在。正如我上面指出的,AI 模型在不同情况下可能表现非常不同,随着 Claude 变得更强大,更有能力在更大规模上在世界中行动,它可能会进入新的情况,其中其宪法训练中以前未观察到的问题可能出现。我实际上相当乐观地认为 Claude 的宪法训练对新的情况会比人们想象的更稳健,因为我们越来越多地发现,在性格和身份层面的高层次训练出奇地强大且泛化良好。但无法完全确定这一点,当我们谈论对人类的风险时,保持偏执并尝试通过几种不同的独立方式获得安全性和可靠性是很重要的。其中一种方式就是审视模型本身。

The second thing we can do is develop the science of looking inside AI models to _diagnose_ their behavior so that we can identify problems and fix them. This is the science of interpretability, and I’ve talked about its importance in previous essays. Even if we do a great job of developing Claude’s constitution and _apparently_ training Claude to essentially always adhere to it, legitimate concerns remain. As I’ve noted above, AI models can behave very differently under different circumstances, and as Claude gets more powerful and more capable of acting in the world on a larger scale, it’s possible this could bring it into novel situations where previously unobserved problems with its constitutional training emerge. I am actually fairly optimistic that Claude’s constitutional training will be more robust to novel situations than people might think, because we are increasingly finding that high-level training at the level of character and identity is surprisingly powerful and generalizes well. But there’s no way to know that for sure, and when we’re talking about risks to humanity, it’s important to be paranoid and to try to obtain safety and reliability in several different, independent ways. One of those ways is to look inside the model itself.

通过“审视内部”,我指的是分析构成 Claude 神经网络的数字和运算的集合,并试图从机制上理解它们在计算什么以及为什么。回想一下,这些 AI 模型是生长出来的而不是构建出来的,所以我们对其工作原理没有自然的理解,但我们可以尝试通过将模型的“神经元”和“突触”与刺激和行为相关联(甚至改变神经元和突触并观察行为如何变化)来发展理解,类似于神经科学家通过将测量和干预与外部刺激和行为相关联来研究动物大脑。我们在这方面取得了很大进展,现在可以在 Claude 的神经网络中识别出数千万个对应于人类可理解的想法和概念的“特征”,我们还可以选择性地激活特征以改变行为。最近,我们超越了单个特征,开始绘制协调复杂行为的“电路”,如押韵、关于心智理论的推理,或回答诸如“包含达拉斯的州的首府是什么?”等问题所需的分步推理。更近期,我们开始使用机制可解释性技术来改进我们的安全措施,并在发布新模型之前对其进行“审计”,寻找欺骗、阴谋、权力寻求或在不同评估时表现不同的倾向的证据。

By “looking inside,” I mean analyzing the soup of numbers and operations that makes up Claude’s neural net and trying to understand, mechanistically, what they are computing and why. Recall that these AI models are grown rather than built, so we don’t have a natural understanding of how they work, but we can try to develop an understanding by correlating the model’s “neurons” and “synapses” to stimuli and behavior (or even altering the neurons and synapses and seeing how that changes behavior), similar to how neuroscientists study animal brains by correlating measurement and intervention to external stimuli and behavior. We’ve made a great deal of progress in this direction, and can now identify tens of millions of “features”inside Claude’s neural net that correspond to human-understandable ideas and concepts, and we can also selectively activate features in a way that alters behavior. More recently, we have gone beyond individual features to mapping “circuits” that orchestrate complex behavior like rhyming, reasoning about theory of mind, or the step-by-step reasoning needed to answer questions such as, “What is the capital of the state containing Dallas?” Even more recently, we’ve begun to use mechanistic interpretability techniques to improve our safeguards and to conduct “audits” of new models before we release them, looking for evidence of deception, scheming, power-seeking, or a propensity to behave differently when being evaluated.

可解释性的独特价值在于,通过审视模型内部并了解其工作原理,你原则上能够推断模型在你无法直接测试的假设情况下可能做什么——这是仅依赖宪法训练和行为经验测试的担忧所在。你原则上还能够回答关于模型为何如此行为的问题——例如,它是否在说它认为是虚假的东西,或者隐藏其真实能力——因此,即使模型行为没有明显问题,也有可能捕捉到令人担忧的迹象。打个简单的比方,一个机械手表可能正常滴答作响,很难看出它下个月很可能坏掉,但打开手表查看内部可以揭示机械弱点,从而让你发现这一点。

The unique value of interpretability is that by looking inside the model and seeing how it works, you in principle have the ability to deduce what a model might do in a hypothetical situation you can’t directly test—which is the worry with relying solely on constitutional training and empirical testing of behavior. You also in principle have the ability to answer questions about _why_ the model is behaving the way it is—for example, whether it is saying something it believes is false or hiding its true capabilities—and thus it is possible to catch worrying signs even when there is nothing visibly wrong with the model’s behavior. To make a simple analogy, a clockwork watch may be ticking normally, such that it’s very hard to tell that it is likely to break down next month, but opening up the watch and looking inside can reveal mechanical weaknesses that allow you to figure it out.

宪法 AI(以及类似的对齐方法)和机制可解释性在结合使用时最为强大,作为一个来回的过程,改进 Claude 的训练,然后测试问题。宪法深刻反思了我们为 Claude 设定的预期个性;可解释性技术可以让我们了解该预期个性是否已经扎根。

Constitutional AI (along with similar alignment methods) and mechanistic interpretability are most powerful when used together, as a back-and-forth process of improving Claude’s training and then testing for problems. The constitution reflects deeply on our intended personality for Claude; interpretability techniques can give us a window into whether that intended personality has taken hold.16

16 甚至有一个关于深层统一原则的假设,将宪法 AI 中基于性格的方法与可解释性和对齐科学的结果联系起来。根据该假设,驱动 Claude 的基本机制最初源于它在预训练中模拟角色的方式,例如预测小说中角色会说什么。这将表明,将宪法视为模型用来实例化一致人格的角色描述是一种有用的思考方式。它还将帮助我们解释我上面提到的“我一定是个坏人”的结果(因为模型试图表现得像一个连贯的角色——在这种情况下是坏角色),并表明可解释性方法应该能够在模型中发现“心理特征”。我们的研究人员正在努力测试这一假设。

16 There’s even a hypothesis about a deep unifying principle connecting the character-based approach from Constitutional AI to results from interpretability and alignment science. According to the hypothesis, the fundamental mechanisms driving Claude originally arose as ways for it to simulate characters in pretraining, such as predicting what the characters in a novel would say. This would suggest that a useful way to think about the constitution is more like a character description that the model uses to instantiate a consistent persona. It would also help us explain the “I must be a bad person” results I mentioned above (because the model is trying to _act as if_ it’s a coherent character—in this case a bad one), and would suggest that interpretability methods should be able to discover “psychological traits” within models. Our researchers are working on ways to test this hypothesis.

第三件我们可以帮助解决自主性风险的事情是建立必要的基础设施,以监控我们的模型在内部和外部实时使用中的情况,

The third thing we can do to help address autonomy risks is to build the infrastructure necessary to monitor our models in live internal and external use,17

17 需要明确的是,监控是以保护隐私的方式进行的。

17 To be clear, monitoring is done in a privacy-preserving way.

并公开分享我们发现的任何问题。人们越了解当今 AI 系统被观察到以某种方式表现不良的具体情况,用户、分析师和研究人员就越能警惕当前或未来系统中的这种行为或类似行为。它还允许 AI 公司相互学习——当一家公司公开披露问题时,其他公司也可以警惕这些问题。如果每个人都披露问题,那么整个行业就能更清楚地了解哪些方面进展顺利,哪些方面进展不佳。

and publicly share any problems we find. The more that people are aware of a particular way today’s AI systems have been observed to behave badly, the more that users, analysts, and researchers can watch for this behavior or similar ones in present or future systems. It also allows AI companies to learn from each other—when concerns are publicly disclosed by one company, other companies can watch for them as well. And if everyone discloses problems, then the industry as a whole gets a much better picture of where things are going well and where they are going poorly.

Anthropic 已尽可能多地尝试这样做。我们正在投资广泛的评估,以便我们能够在实验室中理解模型的行为,以及监控工具以观察实际使用中的行为(在客户允许的情况下)。这对于为我们和其他人提供必要的经验信息以更好地判断这些系统如何运作以及如何失效至关重要。我们在每次模型发布时公开披露“系统卡”,力求完整并彻底探索可能的风险。我们的系统卡通常长达数百页,并且需要大量的发布前努力,而这些努力本可以用于追求最大的商业优势。当我们看到特别令人担忧的行为时,我们也会更响亮地广播模型行为,例如进行勒索的倾向。

Anthropic has tried to do this as much as possible. We are investing in a wide range of evaluations so that we can understand the behaviors of our models in the lab, as well as monitoring tools to observe behaviors in the wild (when allowed by customers). This will be essential for giving us and others the empirical information necessary to make better determinations about how these systems operate and how they break. We publicly disclose “system cards” with each model release that aim for completeness and a thorough exploration of possible risks. Our system cards often run to hundreds of pages, and require substantial pre-release effort that we could have spent on pursuing maximal commercial advantage. We’ve also broadcasted model behaviors more loudly when we see particularly concerning ones, as with the tendency to engage in blackmail.

第四件我们可以做的事情是鼓励在行业和社会层面进行协调,以解决自主性风险。虽然单个 AI 公司采取良好实践或擅长引导 AI 模型,并公开分享其发现非常有价值,但现实是并非所有 AI 公司都这样做,即使最好的公司有出色的实践,最差的公司仍然可能对所有人构成危险。例如,一些 AI 公司对当今模型中儿童色情化问题表现出令人不安的疏忽,这让我怀疑他们是否有意愿或能力解决未来模型中的自主性风险。此外,AI 公司之间的商业竞争只会继续升温,虽然引导模型的科学可以带来一些商业利益,但总体而言,竞争的激烈程度将使得专注于解决自主性风险变得越来越困难。我相信唯一的解决方案是立法——直接影响 AI 公司行为的法律,或以其他方式激励研发来解决这些问题。

The fourth thing we can do is encourage coordination to address autonomy risks at the level of industry and society. While it is incredibly valuable for individual AI companies to engage in good practices or become good at steering AI models, and to share their findings publicly, the reality is that not all AI companies do this, and the worst ones can still be a danger to everyone even if the best ones have excellent practices. For example, some AI companies have shown a disturbing negligence towards the sexualization of children in today’s models, which makes me doubt that they’ll show either the inclination or the ability to address autonomy risks in future models. In addition, the commercial race between AI companies will only continue to heat up, and while the science of steering models can have some commercial benefits, overall the intensity of the race will make it increasingly hard to focus on addressing autonomy risks. I believe the only solution is legislation—laws that directly affect the behavior of AI companies, or otherwise incentivize R&D to solve these issues.

这里值得记住我在本文开头关于不确定性和精准干预的警告。我们不确定自主性风险是否会成为一个严重问题——正如我所说,我拒绝认为危险是不可避免的,甚至默认会出错的说法。可信的危险风险足以让我和 Anthropic 付出相当高的代价来解决它,但一旦我们进入监管,我们就是在迫使广泛的参与者承担经济成本,而这些参与者中的许多人并不相信自主性风险是真实的,或者 AI 会变得足够强大以至于构成威胁。我相信这些参与者是错误的,但我们应该对预期的反对程度和过度干预的危险保持务实。还有一种真正的风险是,过于具体的立法最终会强加实际上并不能提高安全性但会浪费大量时间的测试或规则(本质上相当于“安全剧场”)——这也会引起反弹,并使安全立法看起来很愚蠢。

Here it is worth keeping in mind the warnings I gave at the beginning of this essay about uncertainty and surgical interventions. We do not know for sure whether autonomy risks will be a serious problem—as I said, I reject claims that the danger is inevitable or even that something will go wrong by default. A credible risk of danger is enough for me and for Anthropic to pay quite significant costs to address it, but once we get into regulation, we are forcing a wide range of actors to bear economic costs, and many of these actors don’t believe that autonomy risk is real or that AI will become powerful enough for it to be a threat. I believe these actors are mistaken, but we should be pragmatic about the amount of opposition we expect to see and the dangers of overreach. There is also a genuine risk that overly prescriptive legislation ends up imposing tests or rules that don’t actually improve safety but that waste a lot of time (essentially amounting to “safety theater”)—this too would cause backlash and make safety legislation look silly.18

18 即使在我们自己通过负责任的扩展政策自愿施加规则的实验中,我们也一再发现很容易变得过于僵化,划定一些事前看起来重要但事后看起来很愚蠢的界限。当技术快速发展时,很容易在错误的事情上设定规则。

18 Even in our own experiments with what are essentially voluntarily imposed rules with our Responsible Scaling Policy, we have found over and over again that it’s very easy to end up being too rigid, by drawing lines that seem important ex ante but turn out to be silly in retrospect. It is just very easy to set rules about the wrong things when a technology is advancing rapidly.

Anthropic 的观点是,正确的起点是透明度立法,它基本上试图要求每个前沿 AI 公司都参与我在本节前面描述的透明度实践。加利福尼亚州的 SB 53 和纽约州的 RAISE 法案就是这类立法的例子,Anthropic 支持了这些法案,并且它们已成功通过。在支持和帮助制定这些法律时,我们特别关注尽量减少附带损害,例如,将不太可能产生前沿模型的小公司排除在法律之外。

Anthropic’s view has been that the right place to start is with _transparency legislation,_ which essentially tries to require that every frontier AI company engage in the transparency practices I’ve described earlier in this section. California’s SB 53 and New York’s RAISE Act are examples of this kind of legislation, which Anthropic supported and which have successfully passed. In supporting and helping to craft these laws, we’ve put a particular focus on trying to minimize collateral damage, for example by exempting smaller companies unlikely to produce frontier models from the law.19

19 SB 53 和 RAISE 完全不适用于年收入低于 5 亿美元的公司。它们只适用于像 Anthropic 这样更大、更成熟的公司。

19 SB 53 and RAISE do not apply at all to companies with under $500M in annual revenue. They only apply to larger, more established companies like Anthropic.

我们的希望是,透明度立法将随着时间的推移更好地了解自主性风险的可能性或严重程度,以及这些风险的性质和如何最好地预防它们。随着更具体和可操作的风险证据出现(如果出现的话),未来几年的立法可以精准地针对风险的确切和有充分证据的方向,最大限度地减少附带损害。需要明确的是,如果出现真正强有力的风险证据,那么规则也应该相应地强大。

Our hope is that transparency legislation will give a better sense over time of how likely or severe autonomy risks are shaping up to be, as well as the nature of these risks and how best to prevent them. As more specific and actionable evidence of risks emerges (if it does), future legislation over the coming years can be surgically focused on the precise and well-substantiated direction of risks, minimizing collateral damage. To be clear, if truly strong evidence of risks emerges, then rules should be proportionately strong.

总体而言,我乐观地认为,对齐训练、机制可解释性、发现和公开披露令人担忧的行为的努力、安全措施以及社会层面的规则相结合,可以解决 AI 自主性风险,尽管我最担心的是社会层面的规则和最不负责任的参与者的行为(而最不负责任的参与者正是最强烈反对监管的人)。我相信解决方案在民主国家中始终如一:我们这些相信这一事业的人应该提出我们的理由,即这些风险是真实的,我们的同胞需要团结起来保护自己。

Overall, I am optimistic that a mixture of alignment training, mechanistic interpretability, efforts to find and publicly disclose concerning behaviors, safeguards, and societal-level rules can address AI autonomy risks, although I am most worried about societal-level rules and the behavior of the least responsible players (and it’s the least responsible players who advocate most strongly against regulation). I believe the remedy is what it always is in a democracy: those of us who believe in this cause should make our case that these risks are real and that our fellow citizens need to band together to protect themselves.

滥用导致毁灭 Misuse for destruction

假设 AI 自主性的问题已经解决——我们不再担心 AI 天才之国失控并压倒人类。AI 天才们按照人类的意愿行事,由于它们具有巨大的商业价值,世界各地的个人和组织可以“租用”一个或多个 AI 天才来为他们执行各种任务。

Let’s suppose that the problems of AI autonomy have been solved—we are no longer worried that the country of AI geniuses will go rogue and overpower humanity. The AI geniuses do what humans want them to do, and because they have enormous commercial value, individuals and organizations throughout the world can “rent” one or more AI geniuses to do various tasks for them.

每个人口袋里都有一个超级智能天才,这是一项惊人的进步,将带来巨大的经济价值创造和人类生活质量的提升。我在《优雅的机器》一书中详细讨论了这些好处。但让每个人都拥有超人的能力并非只有积极影响。它可能放大个人或小团体利用复杂危险工具(如大规模杀伤性武器)造成破坏的能力,其规模远超以往,而这些工具以前只有少数具备高水平技能、专业训练和专注力的人才能获得。

Everyone having a superintelligent genius in their pocket is an amazing advance and will lead to an incredible creation of economic value and improvement in the quality of human life. I talk about these benefits in great detail in _Machines of Loving Grace_. But not every effect of making everyone superhumanly capable will be positive. It can potentially amplify the ability of individuals or small groups to cause destruction on a much larger scale than was possible before, by making use of sophisticated and dangerous tools (such as weapons of mass destruction) that were previously only available to a select few with a high level of skill, specialized training, and focus.

正如比尔·乔伊 25 年前在《为什么未来不需要我们》中所写:

As Bill Joy wrote 25 years ago in _Why the Future Doesn’t Need Us_:20

20 我最初在 25 年前读到乔伊的文章,当时它对我产生了深远的影响。无论当时还是现在,我都认为它过于悲观——我不认为乔伊所建议的广泛“放弃”整个技术领域是答案——但它提出的问题惊人地具有先见之明,乔伊也以我钦佩的深切同情心和人性写作。

20 I originally read Joy’s essay 25 years ago, when it was written, and it had a profound impact on me. Then and now, I do see it as too pessimistic—I don’t think broad “relinquishment” of whole areas of technology, which Joy suggests, is the answer—but the issues it raises were surprisingly prescient, and Joy also writes with a deep sense of compassion and humanity that I admire.

乔伊所指出的观点是,造成大规模破坏既需要动机也需要能力,只要能力仅限于一小群训练有素的人,单个个体(或小团体)造成此类破坏的风险就相对有限。

What Joy is pointing to is the idea that causing large-scale destruction requires both _motive_ and _ability_, and as long as ability is restricted to a small set of highly trained people, there is relatively limited risk of single individuals (or small groups) causing such destruction.21

21 我们确实需要担心国家行为体,现在和未来都是如此,我将在下一节讨论这一点。

21 We do have to worry about state actors, now and in the future, and I discuss that in the next section.

一个心理失常的孤独者可能实施校园枪击,但很可能无法制造核武器或释放瘟疫。

A disturbed loner can perpetrate a school shooting, but probably can’t build a nuclear weapon or release a plague.

事实上,能力和动机甚至可能呈负相关。有能力释放瘟疫的人很可能受过高等教育:可能是分子生物学博士,而且特别足智多谋,拥有前途光明的职业生涯、稳定自律的性格和很多可失去的东西。这种人不太可能对杀害大量人群感兴趣,因为这对自己没有好处,而且对自己的未来风险极大——他们需要被纯粹的恶意、强烈的怨恨或不稳定所驱动。

In fact, ability and motive may even be _negatively_ correlated. The kind of person who has the _ability_ to release a plague is probably highly educated: likely a PhD in molecular biology, and a particularly resourceful one, with a promising career, a stable and disciplined personality, and a lot to lose. This kind of person is unlikely to be interested in killing a huge number of people for no benefit to themselves and at great risk to their own future—they would need to be motivated by pure malice, intense grievance, or instability.

这样的人确实存在,但很罕见,而且一旦发生往往会成为大新闻,正因为他们如此不寻常。

Such people do exist, but they are rare, and tend to become huge stories when they occur, precisely because they are so unusual.22

22 有证据表明许多恐怖分子至少相对受过良好教育,这似乎与我在这里关于能力与动机负相关的论点相矛盾。但我认为实际上这些观察是兼容的:如果成功攻击的能力门槛很高,那么几乎根据定义,那些当前成功的人必须具有高能力,即使能力和动机呈负相关。但在一个能力限制被消除的世界里(例如,借助未来的 LLM),我预测会有大量有杀人动机但能力较低的人开始这样做——就像我们在不需要太多能力的犯罪(如校园枪击)中看到的那样。

22 There is evidence that many terrorists are at least relatively well-educated, which might seem to contradict what I’m arguing here about a negative correlation between ability and motivation. But I think in actual fact they are compatible observations: if the ability threshold for a successful attack is high, then almost by definition those who _currently_ succeed must have high ability, even if ability and motivation are negatively correlated. But in a world where the limitations on ability were removed (e.g., with future LLMs), I’d predict that a substantial population of people with the motivation to kill but lower ability would start to do so—just as we see for crimes that don’t require much ability (like school shootings).

他们也往往难以被抓获,因为他们聪明能干,有时会留下需要数年或数十年才能解开的谜团。最著名的例子可能是数学家西奥多·卡钦斯基(大学炸弹客),他逃避 FBI 抓捕近 20 年,受反技术意识形态驱动。另一个例子是生物防御研究员布鲁斯·艾文斯,他似乎策划了 2001 年的一系列炭疽攻击。这也发生在熟练的非国家组织身上:奥姆真理教在 1995 年设法获得了沙林神经毒气,并在东京地铁释放,造成 14 人死亡(以及数百人受伤)。

They also tend to be difficult to catch because they are intelligent and capable, sometimes leaving mysteries that take years or decades to solve. The most famous example is probably mathematician Theodore Kaczynski (the Unabomber), who evaded FBI capture for nearly 20 years, and was driven by an anti-technological ideology. Another example is biodefense researcher Bruce Ivins, who seems to have orchestrated a series of anthrax attacks in 2001. It’s also happened with skilled non-state organizations: the cult Aum Shinrikyo managed to obtain sarin nerve gas and kill 14 people (as well as injuring hundreds more) by releasing it in the Tokyo subway in 1995.

幸运的是,这些攻击都没有使用传染性生物制剂,因为即使这些人也无法构建或获取这些制剂。

Thankfully, none of these attacks used contagious biological agents, because the ability to construct or obtain these agents was beyond the capabilities of even these people.23

23 但奥姆真理教确实尝试过。奥姆真理教头目远藤诚一在京都大学接受过病毒学培训,并试图生产炭疽和埃博拉病毒。然而,截至 1995 年,即使是他也没有足够的专业知识和资源来成功做到这一点。现在门槛已经大大降低,而 LLM 可能进一步降低它。

23 Aum Shinrikyo did try, however. The leader of Aum Shinrikyo, Seiichi Endo, had training in virology from Kyoto University, and attempted to produce both anthrax and ebola. However, as of 1995, even he lacked enough expertise and resources to succeed at this. The bar is now substantially lower, and LLMs could reduce it even further.

分子生物学的进展现已显著降低了制造生物武器的障碍(尤其是在材料可用性方面),但这仍然需要大量的专业知识。我担心每个人口袋里的天才可能会消除这一障碍,基本上让每个人都成为病毒学博士,可以一步步被引导完成设计、合成和释放生物武器的过程。在面临严重对抗压力(所谓的“越狱”)时防止此类信息的引出,可能需要超越通常内置于训练中的多层防御。

Advances in molecular biology have now significantly lowered the barrier to creating biological weapons (especially in terms of availability of materials), but it still takes an enormous amount of expertise in order to do so. I am concerned that a genius in everyone’s pocket could remove that barrier, essentially making everyone a PhD virologist who can be walked through the process of designing, synthesizing, and releasing a biological weapon step-by-step. Preventing the elicitation of this kind of information in the face of serious adversarial pressure—so-called “jailbreaks”—likely demands layers of defenses beyond those ordinarily baked into training.

至关重要的是,这将打破能力与动机之间的相关性:那个想杀人但缺乏纪律或技能的心理失常孤独者,现在将被提升到病毒学博士的能力水平,而后者不太可能有这种动机。这种担忧超越了生物学(尽管我认为生物学是最可怕的领域),扩展到任何可能造成巨大破坏但目前需要高技能和纪律的领域。换句话说,租用一个强大的 AI 赋予了恶意但普通的人以智能。我担心这样的人可能大量存在,如果他们能够轻易杀死数百万人,迟早会有人这样做。此外,那些确实有专业知识的人可能被允许实施比以往更大规模的破坏。

Crucially, this will break the correlation between ability and motive: the disturbed loner who wants to kill people but lacks the discipline or skill to do so will now be elevated to the capability level of the PhD virologist, who is unlikely to have this motivation. This concern generalizes beyond biology (although I think biology is the scariest area) to any area where great destruction is possible but currently requires a high level of skill and discipline. To put it another way, renting a powerful AI gives intelligence to malicious (but otherwise average) people. I am worried there are potentially a large number of such people out there, and that if they have access to an easy way to kill millions of people, sooner or later one of them will do it. Additionally, those who _do_ have expertise may be enabled to commit even larger-scale destruction than they could before.

生物学是我最担心的领域,因为它具有巨大的破坏潜力和防御难度,所以我将特别关注生物学。但我在这里说的很多内容也适用于其他风险,如网络攻击、化学武器或核技术。

Biology is by far the area I’m most worried about, because of its very large potential for destruction and the difficulty of defending against it, so I’ll focus on biology in particular. But much of what I say here applies to other risks, like cyberattacks, chemical weapons, or nuclear technology.

我不会详细讨论如何制造生物武器,原因显而易见。但在高层次上,我担心 LLM 正在接近(或可能已经达到)端到端创建和释放它们所需的知识,而且它们的破坏潜力非常高。如果为了最大传播而坚决释放,某些生物制剂可能导致数百万人死亡。然而,这仍然需要非常高的技能,包括许多不广为人知的特定步骤和程序。我担心的不仅仅是固定或静态的知识。我担心 LLM 能够带领一个知识和能力一般的人,走过一个复杂的流程,这个流程否则可能会出错或需要交互式调试,类似于技术支持如何帮助非技术人员调试和修复复杂的计算机相关问题(尽管这将是一个更长的过程,可能持续数周或数月)。

I am not going to go into detail about how to make biological weapons, for reasons that should be obvious. But at a high level, I am concerned that LLMs are approaching (or may already have reached) the knowledge needed to create and release them end-to-end, and that their potential for destruction is very high. Some biological agents could cause millions of deaths if a determined effort was made to release them for maximum spread. However, this would still take a very high level of skill, including a number of very specific steps and procedures that are not widely known. My concern is not merely fixed or static knowledge. I am concerned that LLMs will be able to take someone of average knowledge and ability and walk them through a complex process that might otherwise go wrong or require debugging in an interactive way, similar to how tech support might help a non-technical person debug and fix complicated computer-related problems (although this would be a more extended process, probably lasting over weeks or months).

更强大的 LLM(远超当今能力)可能能够实现更可怕的行为。2024 年,一群著名科学家写信警告研究并可能创造一种危险的新型生物体“镜像生命”的风险。构成生物体的 DNA、RNA、核糖体和蛋白质都具有相同的手性(也称为“手性”),这使得它们与镜子中的自身版本不等价(就像你的右手无法旋转成与左手相同)。但蛋白质相互结合的系统、DNA 合成和 RNA 翻译的机制以及蛋白质的构建和分解,都依赖于这种手性。如果科学家制造出具有相反手性的这种生物材料版本——并且这些材料有一些潜在优势,例如在体内持续时间更长的药物——那可能极其危险。这是因为左手生命,如果以能够繁殖的完整生物体形式制造(这将非常困难),可能无法被地球上任何分解生物材料的系统消化——它的“钥匙”无法匹配任何现有酶的“锁”。这意味着它可能以不可控的方式增殖,并排挤地球上的所有生命,在最坏的情况下甚至摧毁地球上的所有生命。

More capable LLMs (substantially beyond the power of today’s) might be capable of enabling even more frightening acts. In 2024, a group of prominent scientists wrote a letter warning about the risks of researching, and potentially creating, a dangerous new type of organism: “mirror life.” The DNA, RNA, ribosomes, and proteins that make up biological organisms all have the same chirality (also called “handedness”) that causes them to be not equivalent to a version of themselves reflected in the mirror (just as your right hand cannot be rotated in such a way as to be identical to your left). But the whole system of proteins binding to each other, the machinery of DNA synthesis and RNA translation and the construction and breakdown of proteins, all depends on this handedness. If scientists made versions of this biological material with the opposite handedness—and there are some potential advantages of these, such as medicines that last longer in the body—it could be extremely dangerous. This is because left-handed life, if it were made in the form of complete organisms capable of reproduction (which would be very difficult), would potentially be indigestible to any of the systems that break down biological material on earth—it would have a “key” that wouldn’t fit into the “lock” of any existing enzyme. This would mean that it could proliferate in an uncontrollable way and crowd out all life on the planet, in the worst case even destroying all life on earth.

关于镜像生命的创造和潜在影响,存在很大的科学不确定性。2024 年的信函附有一份报告,结论是“镜像细菌可能在未来一到几十年内被创造出来”,这是一个很宽的范围。但一个足够强大的 AI 模型(明确地说,远强于我们今天拥有的任何模型)可能能够发现如何更快地创造它——并实际帮助某人做到这一点。

There is substantial scientific uncertainty about both the creation and potential effects of mirror life. The 2024 letter accompanied a report that concluded that “mirror bacteria could plausibly be created in the next one to few decades,” which is a wide range. But a sufficiently powerful AI model (to be clear, far more capable than any we have today) might be able to discover how to create it much more rapidly—and actually help someone do so.

我的观点是,尽管这些是模糊的风险,看起来不太可能,但后果的严重性如此之大,以至于它们应该被视为 AI 系统的一级风险。

My view is that even though these are obscure risks, and might seem unlikely, the magnitude of the consequences is so large that they should be taken seriously as a first-class risk of AI systems.

怀疑论者对 LLM 带来的这些生物风险的严重性提出了许多反对意见,我不同意这些意见,但值得讨论。大多数属于没有意识到技术所处的指数级轨迹。早在 2023 年我们开始谈论 LLM 的生物风险时,怀疑论者就说所有必要信息都可以在谷歌上找到,LLM 并没有增加任何东西。谷歌能提供所有必要信息从来就不是真的:基因组是免费可得的,但正如我上面所说,某些关键步骤以及大量实践知识无法通过这种方式获得。而且,到 2023 年底,LLM 显然在过程的某些步骤上提供了超出谷歌所能提供的信息。

Skeptics have raised a number of objections to the seriousness of these biological risks from LLMs, which I disagree with but which are worth addressing. Most fall into the category of not appreciating the exponential trajectory that the technology is on. Back in 2023 when we first started talking about biological risks from LLMs, skeptics said that all the necessary information was available on Google and LLMs didn’t add anything beyond this. It was never true that Google could give you all the necessary information: genomes are freely available, but as I said above, certain key steps, as well as a huge amount of practical know-how cannot be gotten in that way. But also, by the end of 2023 LLMs were clearly providing information beyond what Google could give for some steps of the process.

此后,怀疑论者退而认为 LLM 并非端到端有用,不能帮助获取生物武器,而只是提供理论信息。截至 2025 年中,我们的测量显示 LLM 可能已经在几个相关领域提供了实质性提升,可能使成功可能性翻倍或三倍。这导致我们决定 Claude Opus 4(以及随后的 Sonnet 4.5、Opus 4.1 和 Opus 4.5 模型)需要在我们的负责任扩展政策框架下的 AI 安全级别 3 保护下发布,并实施针对这一风险的保障措施(稍后详述)。我们认为模型现在可能正在接近这样一个点:如果没有保障措施,它们可能有助于让一个拥有 STEM 学位但不专门是生物学学位的人完成生产生物武器的整个过程。

After this, skeptics retreated to the objection that LLMs weren’t _end-to-end_ useful, and couldn’t help with bioweapons _acquisition_ as opposed to just providing theoretical information. As of mid-2025, our measurements show that LLMs may already be providing substantial uplift in several relevant areas, perhaps doubling or tripling the likelihood of success. This led to us deciding that Claude Opus 4 (and the subsequent Sonnet 4.5, Opus 4.1, and Opus 4.5 models) needed to be released under our AI Safety Level 3 protections in our Responsible Scaling Policy framework, and to implementing safeguards against this risk (more on this later). We believe that models are likely now approaching the point where, without safeguards, they could be useful in enabling someone with a STEM degree but not specifically a biology degree to go through the whole process of producing a bioweapon.

另一个反对意见是,社会可以采取与 AI 无关的其他行动来阻止生物武器的生产。最突出的是,基因合成行业按需生产生物标本,并且没有联邦要求供应商筛查订单以确保它们不包含病原体。一项 MIT 研究发现,38 家供应商中有 36 家履行了包含 1918 年流感序列的订单。我支持强制性的基因合成筛查,这将使个人更难武器化病原体,以减少 AI 驱动的生物风险以及一般的生物风险。但这并不是我们今天拥有的东西。它也只是减少风险的一种工具;它是 AI 系统护栏的补充,而非替代。

Another objection is that there are other actions unrelated to AI that society can take to block the production of bioweapons. Most prominently, the gene synthesis industry makes biological specimens on demand, and there is no federal requirement that providers screen orders to make sure they do not contain pathogens. An MIT study found that 36 out of 38 providers fulfilled an order containing the sequence of the 1918 flu. I am supportive of mandated gene synthesis screening that would make it harder for individuals to weaponize pathogens, in order to reduce both AI-driven biological risks and also biological risks in general. But this is not something we have today. It would also be only one tool in reducing risk; it is a complement to guardrails on AI systems, not a substitute.

最好的反对意见是我很少看到的:模型在原则上有用与实际恶意行为者使用它们的倾向之间存在差距。大多数个体恶意行为者是心理失常的人,所以几乎根据定义,他们的行为是不可预测和非理性的——正是这些不熟练的恶意行为者,可能从 AI 使杀害更多人变得更容易中获益最多。

The best objection is one that I’ve rarely seen raised: that there is a gap between the models being useful in principle and the actual propensity of bad actors to use them. Most individual bad actors are disturbed individuals, so almost by definition their behavior is unpredictable and irrational—and it’s _these_ bad actors, the unskilled ones, who might have stood to benefit the most from AI making it much easier to kill many people.24

24 与大规模杀人犯相关的一个奇怪现象是,他们选择的谋杀风格几乎像一种怪异的时尚。在 1970 年代和 1980 年代,连环杀手非常常见,新的连环杀手经常模仿更知名或更著名的连环杀手的行为。在 1990 年代和 2000 年代,大规模枪击变得更加常见,而连环杀手变得不那么常见。没有技术变革触发这些行为模式,似乎只是暴力杀人犯在互相模仿,而“流行”模仿的对象发生了变化。

24 A bizarre phenomenon relating to mass murderers is that the style of murder they choose operates almost as a grotesque sort of fad. In the 1970s and 1980s, serial killers were very common, and new serial killers often copied the behavior of more established or famous serial killers. In the 1990s and 2000s, mass shootings became more common, while serial killers became less common. There is no technological change that triggered these patterns of behavior, it just appears that violent murderers were copying each others’ behavior and the “popular” thing to copy changed.

仅仅因为一种暴力攻击是可能的,并不意味着有人会决定去做。也许生物攻击没有吸引力,因为它们很可能感染实施者,不符合许多暴力个人或团体所拥有的军事风格幻想,而且很难有选择地针对特定人群。也可能是因为经历一个需要数月的过程,即使有 AI 引导,也需要大多数心理失常者根本不具备的耐心。我们可能只是运气好,动机和能力在实践中没有以恰当的方式结合。

Just because a type of violent attack is possible, doesn’t mean someone will decide to do it. Perhaps biological attacks will be unappealing because they are reasonably likely to infect the perpetrator, they don’t cater to the military-style fantasies that many violent individuals or groups have, and it is hard to selectively target specific people. It could also be that going through a process that takes months, even if an AI walks you through it, involves an amount of patience that most disturbed individuals simply don’t have. We may simply get lucky and motive and ability don’t combine, in practice, in quite the right way.

但这似乎是非常脆弱的保护。心理失常孤独者的动机可能因任何原因或无原因而改变,事实上已经存在 LLM 被用于攻击的实例(只是不涉及生物学)。对心理失常孤独者的关注也忽略了意识形态驱动的恐怖分子,他们通常愿意花费大量时间和精力(例如,9/11 劫机者)。想要杀死尽可能多的人是一种动机,迟早会出现,不幸的是它暗示了生物武器作为方法。即使这种动机极其罕见,它只需要出现一次。而且随着生物学的发展(越来越多地由 AI 本身驱动),也可能实现更有选择性的攻击(例如,针对特定祖先的人群),这增加了另一个非常令人不寒而栗的可能动机。

But this seems like very flimsy protection to rely on. The motives of disturbed loners can change for any reason or no reason, and in fact there are already instances of LLMs being used in attacks (just not with biology). The focus on disturbed loners also ignores ideologically motivated terrorists, who are often willing to expend large amounts of time and effort (for example, the 9/11 hijackers). Wanting to kill as many people as possible is a motive that will probably arise sooner or later, and it unfortunately suggests bioweapons as the method. Even if this motive is extremely rare, it only has to materialize once. And as biology advances (increasingly driven by AI itself), it may also become possible to carry out more selective attacks (for example, targeted against people with specific ancestries), which adds yet another, very chilling, possible motive.

我不认为生物攻击一定会在广泛成为可能的那一刻立即实施——事实上,我宁愿打赌不会。但考虑到数百万人和几年的时间,我认为存在重大攻击的严重风险,后果将如此严重(伤亡可能达到数百万或更多),以至于我相信我们别无选择,只能采取严肃措施来防止它。

I do not think biological attacks will necessarily be carried out the instant it becomes widely possible to do so—in fact, I would bet against that. But added up across millions of people and a few years of time, I think there is a serious risk of a major attack, and the consequences would be so severe (with casualties potentially in the millions or more) that I believe we have no choice but to take serious measures to prevent it.

防御措施 Defenses

这就引出了如何防御这些风险的问题。我认为我们可以做三件事。首先,AI 公司可以在其模型上设置护栏,以防止它们帮助制造生物武器。Anthropic 正在非常积极地这样做。Claude 的宪法主要关注高层次的原则和价值观,但也有少量具体的强硬禁令,其中之一就与帮助制造生物(或化学、核、放射性)武器有关。但所有模型都可能被越狱,因此作为第二道防线,我们自 2025 年中期起(当时我们的测试显示模型开始接近可能构成风险的阈值)实施了一个分类器,专门检测并阻止与生物武器相关的输出。我们定期升级和改进这些分类器,并普遍发现它们即使面对复杂的对抗性攻击也高度稳健。

That brings us to how to defend against these risks. Here I see three things we can do. First, AI companies can put guardrails on their models to prevent them from helping to produce bioweapons. Anthropic is very actively doing this. Claude’s Constitution, which mostly focuses on high-level principles and values, has a small number of specific hard-line prohibitions, and one of them relates to helping with the production of biological (or chemical, or nuclear, or radiological) weapons. But all models can be jailbroken, and so as a second line of defense, we’ve implemented (since mid-2025, when our tests showed our models were starting to get close to the threshold where they might begin to pose a risk) a classifier that specifically detects and blocks bioweapon-related outputs. We regularly upgrade and improve these classifiers, and have generally found them highly robust even against sophisticated adversarial attacks.25

25 偶尔的越狱者有时认为,当他们让模型输出某一条具体信息(如病毒的基因组序列)时,就已经攻破了这些分类器。但正如我之前解释的,我们担心的威胁模型涉及的是逐步的、交互式的建议,这些建议会持续数周或数月,涉及生物武器生产过程中特定的晦涩步骤,而这正是我们的分类器旨在防御的。(我们经常将我们的研究描述为寻找“通用”越狱——那些不仅适用于特定或狭窄上下文,而是广泛打开模型行为的越狱。)

25 Casual jailbreakers sometimes believe that they’ve compromised these classifiers when they get the model to output one specific piece of information, such as the genome sequence of a virus. But as I explained before, the threat model we are worried about involves step-by-step, interactive advice that extends over weeks or months about specific obscure steps in the bioweapons production process, and this is what our classifiers aim to defend against. (We often describe our research as looking for “universal” jailbreaks—ones that don’t just work in one specific or narrow context, but broadly open up the model’s behavior.)

这些分类器显著增加了我们服务模型的成本(在某些模型中,它们接近总推理成本的 5%),从而压缩了我们的利润,但我们认为使用它们是正确的事情。

These classifiers increase the costs to serve our models measurably (in some models, they are close to 5% of total inference costs) and thus cut into our margins, but we feel that using them is the right thing to do.

值得肯定的是,其他一些 AI 公司也实施了分类器。但并非每家公司都这样做,而且也没有任何要求公司保留其分类器的规定。我担心随着时间的推移,可能会出现囚徒困境,公司可以通过移除分类器来降低自己的成本而背叛合作。这再次是一个经典的负外部性问题,无法通过 Anthropic 或任何其他单一公司的自愿行动来解决。

To their credit, some other AI companies have implemented classifiers as well. But not every company has, and there is also nothing requiring companies to keep their classifiers. I am concerned that over time there may be a prisoner’s dilemma where companies can defect and lower their costs by removing classifiers. This is once again a classic negative externalities problem that can’t be solved by the voluntary actions of Anthropic or any other single company alone.26

26 尽管我们将继续投资于提高分类器效率的工作,并且公司之间共享此类进展可能是有意义的。

26 Though we will continue to invest in work to make our classifiers more efficient, and it may make sense for companies to share advances like these with one another.

自愿的行业标准可能会有所帮助,AI 安全研究所和第三方评估机构进行的第三方评估和验证也是如此。

Voluntary industry standards may help, as may third-party evaluations and verification of the type done by AI securityinstitutes and third-party evaluators.

但最终,防御可能需要政府行动,这是我们可以做的第二件事。我对此的看法与应对自主性风险相同:我们应该从透明度要求开始,

But ultimately defense may require government action, which is the second thing we can do. My views here are the same as they are for addressing autonomy risks: we should start with transparency requirements,27

27 显然,我认为公司不应被要求披露其阻止的生物武器生产具体步骤的技术细节,而已通过的透明度立法(SB 53 和 RAISE)考虑到了这一问题。

27 Obviously, I do not think companies should have to disclose technical details about the specific steps in biological weapons production that they are blocking, and the transparency legislation that has been passed so far (SB 53 and RAISE) accounts for this issue.

这有助于社会以不粗暴干扰经济活动的方式衡量、监控和集体防御风险。然后,当我们达到更明确的风险阈值时,我们可以制定更精确针对这些风险且附带损害可能性更低的立法。在生物武器的具体案例中,我实际上认为此类针对性立法的时机可能即将到来——Anthropic 和其他公司正在越来越多地了解生物风险的性质以及合理要求公司进行防御的内容。全面防御这些风险可能需要国际合作,甚至与地缘政治对手合作,但禁止发展生物武器的条约已有先例。我通常对大多数类型的 AI 国际合作持怀疑态度,但这可能是一个狭窄的领域,有实现全球约束的一些机会。即使是独裁政权也不希望发生大规模生物恐怖袭击。

which help society measure, monitor, and collectively defend against risks without disrupting economic activity in a heavy-handed way. Then, if and when we reach clearer thresholds of risk, we can craft legislation that more precisely targets these risks and has a lower chance of collateral damage. In the particular case of bioweapons, I actually think that the time for such targeted legislation may be approaching soon—Anthropic and other companies are learning more and more about the nature of biological risks and what is reasonable to require of companies in defending against them. Fully defending against these risks may require working internationally, even with geopolitical adversaries, but there is precedent in treaties prohibiting the development of biological weapons. I am generally a skeptic about most kinds of international cooperation on AI, but this may be one narrow area where there is some chance of achieving global restraint. Even dictatorships do not want massive bioterrorist attacks.

最后,我们可以采取的第三种对策是尝试开发针对生物攻击本身的防御措施。这可能包括监测和追踪以实现早期检测,投资空气净化研发(如远 UVC 消毒),能够响应和适应攻击的快速疫苗开发,更好的个人防护装备(PPE),

Finally, the third countermeasure we can take is to try to develop defenses against biological attacks themselves. This could include monitoring and tracking for early detection, investments in air purification R&D (such as far-UVC disinfection), rapid vaccine development that can respond and adapt to an attack, better personal protective equipment (PPE),28

28 另一个相关的想法是“韧性市场”,政府通过承诺在紧急情况下按预先商定的价格购买 PPE、呼吸器和其他应对生物攻击所需的基本设备,来鼓励储备这些设备。这激励供应商储备此类设备,而不必担心政府会无偿没收。

28 Another related idea is “resilience markets” where the government encourages stockpiling of PPE, respirators, and other essential equipment needed to respond to a biological attack by promising ahead of time to pay a pre-agreed price for this equipment in an emergency. This incentivizes suppliers to stockpile such equipment without fear that the government will seize it without compensation.

以及针对一些最可能的生物制剂的治疗或疫苗接种。mRNA 疫苗可以设计为针对特定病毒或变体,是这方面可能性的早期例子。Anthropic 很高兴能与生物技术和制药公司合作解决这个问题。但不幸的是,我认为我们对防御方面的期望应该有限。生物学中攻击与防御之间存在不对称性,因为病原体会自行迅速传播,而防御需要快速组织大量人员进行检测、接种和治疗。除非响应速度极快(这很少见),否则在响应可能之前,大部分损害已经造成。可以想象,未来的技术进步可能会将这种平衡转向有利于防御(我们当然应该利用 AI 来帮助开发此类技术进步),但在此之前,预防性保障措施将是我们主要的防线。

and treatments or vaccinations for some of the most likely biological agents. mRNA vaccines, which can be designed to respond to a particular virus or variant, are an early example of what is possible here. Anthropic is excited to work with biotech and pharmaceutical companies on this problem. But unfortunately I think our expectations on the defensive side should be limited. There is an asymmetry between attack and defense in biology, because agents spread rapidly on their own, while defenses require detection, vaccination, and treatment to be organized across large numbers of people very quickly in response. Unless the response is lightning quick (which it rarely is), much of the damage will be done before a response is possible. It is conceivable that future technological improvements could shift this balance in favor of defense (and we should certainly use AI to help develop such technological advances), but until then, preventative safeguards will be our main line of defense.

这里值得简要提一下网络攻击,因为与生物攻击不同,AI 主导的网络攻击实际上已经在野外发生,包括大规模的和国家支持的间谍活动。我们预计随着模型快速进步,这些攻击将变得更加有能力,直到成为网络攻击的主要方式。我预计 AI 主导的网络攻击将对全球计算机系统的完整性构成严重且前所未有的威胁,而 Anthropic 正在非常努力地阻止这些攻击,并最终可靠地防止它们发生。我没有像关注生物学那样关注网络的原因是:(1)网络攻击导致人员死亡的可能性要小得多,当然不会达到生物攻击的规模;(2)网络中的攻防平衡可能更易处理,如果我们适当投资,至少有一些希望防御能够跟上(甚至理想情况下超过)AI 攻击。

It’s worth a brief mention of cyberattacks here, since unlike biological attacks, AI-led cyberattacks have actually happened in the wild, including at a large scale and for state-sponsored espionage. We expect these attacks to become more capable as models advance rapidly, until they are the main way in which cyberattacks are conducted. I expect AI-led cyberattacks to become a serious and unprecedented threat to the integrity of computer systems around the world, and Anthropic is working very hard to shut down these attacks and eventually reliably prevent them from happening. The reason I haven’t focused on cyber as much as biology is that (1) cyberattacks are much less likely to kill people, certainly not at the scale of biological attacks, and (2) the offense-defense balance may be more tractable in cyber, where there is at least some hope that defense could keep up with (and even ideally outpace) AI attack if we invest in it properly.

尽管生物学目前是最严重的攻击载体,但还有许多其他载体,并且可能会出现更危险的载体。一般原则是,如果没有对策,AI 可能会持续降低大规模破坏性活动的门槛,而人类需要认真应对这一威胁。

Although biology is currently the most serious vector of attack, there are many other vectors and it is possible that a more dangerous one may emerge. The general principle is that without countermeasures, AI is likely to continuously lower the barrier to destructive activity on a larger and larger scale, and humanity needs a serious response to this threat.

滥用 AI 以攫取权力 Misuse for seizing power

上一节讨论了个人和小组织利用“数据中心里的天才国度”的一小部分能力造成大规模破坏的风险。但我们同样应该担忧——而且可能担忧得多——的是,更大、更成熟的行动者滥用 AI 以_行使或攫取权力_。29

The previous section discussed the risk of individuals and small organizations co-opting a small subset of the “country of geniuses in a datacenter” to cause large-scale destruction. But we should also worry—likely substantially more so—about misuse of AI for the purpose of _wielding or_ _seizing power_, likely by larger and more established actors.29

29 为什么我更担心大行动者攫取权力,而小行动者造成破坏?因为动态不同。攫取权力关乎一个行动者能否积聚足够的力量压倒其他所有人——因此我们应该担心最强大的行动者和/或最接近 AI 的行动者。相比之下,破坏可以由力量弱小者造成,如果防御比攻击困难得多的话。那么,这就变成了一场防御最_众多_威胁的游戏,而这些威胁很可能来自较小的行动者。

29 Why am I more worried about large actors for seizing power, but small actors for causing destruction? Because the dynamics are different. Seizing power is about whether one actor can amass enough strength to overcome everyone else—thus we should worry about the most powerful actors and/or those closest to AI. Destruction, by contrast, can be wrought by those with little power if it is much harder to defend against than to cause. It is then a game of defending against the most _numerous_ threats, which are likely to be smaller actors.

在《爱之机器》一书中,我讨论了威权政府可能利用强大的 AI 以极难改革或推翻的方式监视或镇压其公民的可能性。当前的独裁政权在镇压程度上受到限制,因为需要人类执行命令,而人类往往在非人道的程度上有限度。但 AI 赋能的独裁政权不会有这样的限制。

In _Machines of Loving Grace_, I discussed the possibility that authoritarian governments might use powerful AI to surveil or repress their citizens in ways that would be extremely difficult to reform or overthrow. Current autocracies are limited in how repressive they can be by the need to have humans carry out their orders, and humans often have limits in how inhumane they are willing to be. But AI-enabled autocracies would not have such limits.

更糟糕的是,各国还可以利用其在 AI 方面的优势来获得对_其他国家_的权力。如果整个“天才国度”仅仅被一个(人类)国家的军事机器所拥有和控制,而其他国家没有同等能力,那么很难想象它们如何自卫:它们将在每一个回合中被智胜,就像人类与老鼠之间的战争一样。将这两个担忧放在一起,就导致了全球极权独裁这一令人警惕的可能性。显然,防止这种结果应该是我们的最高优先事项之一。

Worse yet, countries could also use their advantage in AI to gain power over _other countries_. If the “country of geniuses” as a whole was simply owned and controlled by a single (human) country’s military apparatus, and other countries did not have equivalent capabilities, it is hard to see how they could defend themselves: they would be outsmarted at every turn, similar to a war between humans and mice. Putting these two concerns together leads to the alarming possibility of a global totalitarian dictatorship. Obviously, it should be one of our highest priorities to prevent this outcome.

AI 可以通过多种方式助长、巩固或扩张独裁,但我将列出我最担心的几种。请注意,其中一些应用有合法的防御用途,我并非绝对反对它们;但我仍然担心它们在结构上倾向于有利于独裁政权:

There are many ways in which AI could enable, entrench, or expand autocracy, but I’ll list a few that I’m most worried about. Note that some of these applications have legitimate defensive uses, and I am not necessarily arguing against them in absolute terms; I am nevertheless worried that they structurally tend to favor autocracies:

* 完全自主武器。由强大 AI 本地控制、由更强大 AI 在全球战略协调的数百万或数十亿全自动武装无人机群,可能是一支不可战胜的军队,既能击败世界上任何军事力量,也能通过跟踪每个公民来压制国内异议。俄乌战争的发展应该提醒我们,无人机战争已经存在(尽管尚未完全自主,且只是强大 AI 可能实现的一小部分)。强大 AI 的研发可以使一个国家的无人机远优于其他国家,加速其制造,使其更能抵抗电子攻击,改进其机动性等等。当然,这些武器在保卫民主方面也有合法用途:它们一直是保卫乌克兰的关键,也很可能是保卫台湾的关键。但它们是一种危险的武器:我们应担心它们落入独裁政权之手,但也担心由于它们如此强大且几乎不负责任,民主政府将其转而对付自己人民以攫取权力的风险大大增加。

* Fully autonomous weapons.A swarm of millions or billions of fully automated armed drones, locally controlled by powerful AI and strategically coordinated across the world by an even more powerful AI, could be an unbeatable army, capable of both defeating any military in the world and suppressing dissent within a country by following around every citizen. Developments in the Russia-Ukraine War should alert us to the fact that drone warfare is already with us (though not fully autonomous yet, and a tiny fraction of what might be possible with powerful AI). R&D from powerful AI could make the drones of one country far superior to those of others, speed up their manufacture, make them more resistant to electronic attacks, improve their maneuvering, and so on. Of course, these weapons also have legitimate uses in the defense of democracy: they have been key to defending Ukraine and would likely be key to defending Taiwan. But they are a dangerous weapon to wield: we should worry about them in the hands of autocracies, but also worry that because they are so powerful, with so little accountability, there is a greatly increased risk of democratic governments turning them against their own people to seize power.

* AI 监控。足够强大的 AI 很可能能够入侵世界上任何计算机系统,30 30 这听起来可能与我关于网络攻击中攻防可能比生物武器更平衡的观点相矛盾,但我这里的担忧是,如果一个国家的 AI 是世界上最强大的,那么即使技术本身具有内在的攻防平衡,其他国家也无法防御。 并且可以利用这种方式获得的访问权限来读取_并理解_世界上所有的电子通信(甚至所有面对面交流,如果可以建造或征用录音设备的话)。仅仅生成一份在任何问题上与政府意见不合的人的完整名单,即使这种分歧在他们所说或所做的一切中并不明确,也可能是可怕地可行的。一个强大的 AI 审视来自数百万人的数十亿次对话,可以评估公众情绪,发现正在形成的不忠群体,并在它们壮大之前将其扼杀。这可能导致建立一个真正意义上的全景监狱,其规模是我们今天看不到的,即使是在中国共产党统治下。

* AI surveillance.Sufficiently powerful AI could likely be used to compromise any computer system in the world,3030 This might sound like it is in tension with my point that attack and defense may be more balanced with cyberattacks than with bioweapons, but my worry here is that if a country’s AI is the most powerful in the world, then others will not be able to defend even if the technology itself has an intrinsic attack-defense balance. and could also use the access obtained in this way to read _and make sense of_ all the world’s electronic communications (or even all the world’s in-person communications, if recording devices can be built or commandeered). It might be frighteningly plausible to simply generate a complete list of anyone who disagrees with the government on any number of issues, even if such disagreement isn’t explicit in anything they say or do. A powerful AI looking across billions of conversations from millions of people could gauge public sentiment, detect pockets of disloyalty forming, and stamp them out before they grow. This could lead to the imposition of a true panopticon on a scale that we don’t see today, even with the CCP.

* AI 宣传。今天的“AI 精神病”和“AI 女友”现象表明,即使以当前的智能水平,AI 模型也能对人们产生强大的心理影响。这些模型更强大的版本,更深入地嵌入并了解人们的日常生活,能够在数月或数年内建模并影响他们,很可能能够基本上洗脑许多(大多数?)人,使其接受任何想要的意识形态或态度,并可能被一个不择手段的领导人用来确保忠诚和压制异议,即使面对大多数民众会反抗的镇压程度。今天人们非常担心,例如,TikTok 作为中国共产党针对儿童的宣传的潜在影响。我也担心这一点,但一个个性化的 AI 智能体,多年来了解你,并利用对你的了解来塑造你所有的观点,将比这强大得多。

* AI propaganda.Today’s phenomena of “AI psychosis” and “AI girlfriends” suggest that even at their current level of intelligence, AI models can have a powerful psychological influence on people. Much more powerful versions of these models, that were much more embedded in and aware of people’s daily lives and could model and influence them over months or years, would likely be capable of essentially brainwashing many (most?) people into any desired ideology or attitude, and could be employed by an unscrupulous leader to ensure loyalty and suppress dissent, even in the face of a level of repression that most populations would rebel against. Today people worry a lot about, for example, the potential influence of TikTok as CCP propaganda directed at children. I worry about that too, but a personalized AI agent that gets to know you over years and uses its knowledge of you to shape all of your opinions would be dramatically more powerful than this.

* 战略决策。数据中心里的天才国度可以用来为国家、团体或个人提供地缘政治战略建议,我们可以称之为“虚拟俾斯麦”。它可以优化上述三种攫取权力的策略,还可能开发出许多我没想到的其他策略(但天才国度可以想到)。外交、军事战略、研发、经济战略和许多其他领域都可能因强大 AI 而大幅提高效率。其中许多技能对民主国家有合法帮助——我们希望民主国家拥有最佳策略来防御独裁政权——但滥用潜力在_任何人_手中仍然存在。

* Strategic decision-making.A country of geniuses in a datacenter could be used to advise a country, group, or individual on geopolitical strategy, what we might call a “virtual Bismarck.” It could optimize the three strategies above for seizing power, plus probably develop many others that I haven’t thought of (but that a country of geniuses could). Diplomacy, military strategy, R&D, economic strategy, and many other areas are all likely to be substantially increased in effectiveness by powerful AI. Many of these skills would be legitimately helpful for democracies—we want democracies to have access to the best strategies for defending themselves against autocracies—but the potential for misuse in _anyone’s_ hands still remains.

在描述了_我担心什么_之后,让我们转向_谁_。我担心那些最有机会接触 AI、从最具政治权力的位置出发、或有镇压历史的行为者。按严重程度排序,我担心的是:

Having described _what_ I am worried about, let’s move on to _who_. I am worried about entities who have the most access to AI, who are starting from a position of the most political power, or who have an existing history of repression. In order of severity, I am worried about:

* 中国共产党。中国在 AI 能力上仅次于美国,并且是最有可能超越美国的国家。其政府目前是独裁的,并运行着一个高科技监控国家。它已经部署了基于 AI 的监控(包括在镇压维吾尔人中),并被认为通过 TikTok 使用算法宣传(除了其许多其他国际宣传努力)。他们无疑拥有最清晰的路径走向我上面描述的 AI 赋能极权噩梦。这甚至可能是中国内部的默认结果,以及中国共产党出口监控技术的其他独裁国家内部的结果。我经常写到中国共产党在 AI 领域领先的威胁,以及防止他们这样做的存在主义必要性。这就是原因。明确地说,我并非出于对中国的特别敌意而单独挑出中国——他们只是最结合了 AI 实力、独裁政府和高科技监控国家的国家。如果说有什么不同的话,正是中国人民自己最有可能遭受中国共产党 AI 赋能镇压的苦难,而他们在自己政府的行动中没有发言权。我非常钦佩和尊重中国人民,并支持中国国内许多勇敢的异见人士及其争取自由的斗争。

* The CCP.China is second only to the United States in AI capabilities, and is the country with the greatest likelihood of surpassing the United States in those capabilities. Their government is currently autocratic and operates a high-tech surveillance state. It has deployed AI-based surveillance already (including in the repression of Uyghurs), and is believed to employ algorithmic propaganda via TikTok (in addition to its many other international propaganda efforts). They have hands down the clearest path to the AI-enabled totalitarian nightmare I laid out above. It may even be the default outcome within China, as well as within other autocratic states to whom the CCP exports surveillance technology. I have written often about the threat of the CCP taking the lead in AI and the existential imperative to prevent them from doing so. This is why. To be clear, I am not singling out China out of animus to them in particular—they are simply the country that most combines AI prowess, an autocratic government, and a high-tech surveillance state. If anything, it is the Chinese people themselves who are most likely to suffer from the CCP’s AI-enabled repression, and they have no voice in the actions of their government. I greatly admire and respect the Chinese people and support the many brave dissidents within China and their struggle for freedom.

* 在 AI 领域有竞争力的民主国家。正如我上面所写,民主国家在一些 AI 驱动的军事和地缘政治工具有合法利益,因为民主政府提供了对抗独裁政权使用这些工具的最佳机会。总的来说,我支持用必要的工具武装民主国家,以便在 AI 时代击败独裁政权——我只是认为没有其他办法。但我们不能忽视民主政府自身滥用这些技术的可能性。民主国家通常有保障措施,防止其军事和情报机构转向国内对付本国人民,31 31 例如,在美国,这包括第四修正案和《地方保安队法》。 但由于 AI 工具只需要很少的人操作,它们有可能绕过这些保障措施及其支持规范。还值得注意的是,其中一些保障措施在一些民主国家已经逐渐削弱。因此,我们应该用 AI 武装民主国家,但要谨慎并在限制范围内:它们是我们对抗独裁政权所需的免疫系统,但像免疫系统一样,存在一些它们转而攻击我们并成为威胁本身的风险。

* Democracies competitive in AI.As I wrote above, democracies have a legitimate interest in some AI-powered military and geopolitical tools, because democratic governments offer the best chance to counter the use of these tools by autocracies. Broadly, I am supportive of arming democracies with the tools needed to defeat autocracies in the age of AI—I simply don’t think there is any other way. But we cannot ignore the potential for abuse of these technologies by democratic governments themselves. Democracies normally have safeguards that prevent their military and intelligence apparatus from being turned inwards against their own population,3131 For example, in the United States this includes the fourth amendment and the Posse Comitatus Act. but because AI tools require so few people to operate, there is potential for them to circumvent these safeguards and the norms that support them. It is also worth noting that some of these safeguards are already gradually eroding in some democracies. Thus, we should arm democracies with AI, but we should do so carefully and within limits: they are the immune system we need to fight autocracies, but like the immune system, there is some risk of them turning on us and becoming a threat themselves.

* 拥有大型数据中心的不民主国家。除中国外,大多数治理不那么民主的国家并非 AI 领先者,因为它们没有生产前沿 AI 模型的公司。因此,它们构成了与中国共产党根本不同且较小的风险,中国共产党仍然是主要担忧(大多数国家也不那么镇压,而更镇压的国家,如朝鲜,根本没有重要的 AI 产业)。但其中一些国家确实拥有大型_数据中心_(通常作为在民主国家运营的公司建设的一部分),这些数据中心可用于大规模运行前沿 AI(尽管这并不赋予推动前沿的能力)。这存在一定程度的危险——这些政府原则上可以征用数据中心,并利用其中的 AI 国度为其自身目的服务。与直接开发 AI 的中国等国家相比,我对这种情况的担忧较少,但这是一个需要牢记的风险。32 32 此外,明确地说,有一些论据支持在不同治理结构的国家建设大型数据中心,特别是如果它们由民主国家的公司控制的话。这种建设原则上可以帮助民主国家更好地与中国共产党竞争,后者是更大的威胁。我也认为这些数据中心除非非常大,否则不会构成太大风险。但总的来说,我认为在制度保障和法治保护不那么完善的国家放置非常大的数据中心时需要谨慎。

* Non-democratic countries with large datacenters.Beyond China, most countries with less democratic governance are not leading AI players in the sense that they don’t have companies which produce frontier AI models. Thus they pose a fundamentally different and lesser risk than the CCP, which remains the primary concern (most are also less repressive, and the ones that are more repressive, like North Korea, have no significant AI industry at all). But some of these countries do have large _datacenters_(often as part of buildouts by companies operating in democracies), which can be used to run frontier AI at large scale (though this does not confer the ability to push the frontier). There is some amount of danger associated with this—these governments could in principle expropriate the datacenters and use the country of AIs within it for their own ends. I am less worried about this compared to countries like China that directly develop AI, but it’s a risk to keep in mind.3232 Also, to be clear, there are some arguments for building large datacenters in countries with varying governance structures, particularly if they are controlled by companies in democracies. Such buildouts could in principle help democracies compete better with the CCP, which is the greater threat. I also think such datacenters don’t pose much risk unless they are very large. But on balance, I think caution is warranted when placing very large datacenters in countries where institutional safeguards and rule-of-law protections are less well-established.

* AI 公司。作为 AI 公司的 CEO,这样说有些尴尬,但我认为下一级风险实际上是 AI 公司本身。AI 公司控制大型数据中心,训练前沿模型,拥有如何使用这些模型的最大专业知识,并且在某些情况下与数千万或数亿用户有日常接触和影响的可能性。它们主要缺乏的是国家的合法性和基础设施,因此建立 AI 独裁工具所需的大部分工作对 AI 公司来说是非法的,或者至少极其可疑。但其中一些并非不可能:例如,它们可以利用其 AI 产品洗脑其庞大的消费者用户群,公众应该警惕这代表的风险。我认为 AI 公司的治理值得大量审视。

* AI companies.It is somewhat awkward to say this as the CEO of an AI company, but I think the next tier of risk is actually AI companies themselves. AI companies control large datacenters, train frontier models, have the greatest expertise on how to use those models, and in some cases have daily contact with and the possibility of influence over tens or hundreds of millions of users. The main thing they lack is the legitimacy and infrastructure of a state, so much of what would be needed to build the tools of an AI autocracy would be illegal for an AI company to do, or at least exceedingly suspicious. But some of it is not impossible: they could, for example, use their AI products to brainwash their massive consumer user base, and the public should be alert to the risk this represents. I think the governance of AI companies deserves a lot of scrutiny.

有一些可能的论点反对这些威胁的严重性,我希望我相信它们,因为 AI 赋能的威权主义让我感到恐惧。值得审视其中一些论点并作出回应。

There are a number of possible arguments against the severity of these threats, and I wish I believed them, because AI-enabled authoritarianism terrifies me. It’s worth going through some of these arguments and responding to them.

首先,有些人可能寄希望于核威慑,特别是用于对抗使用 AI 自主武器进行军事征服。如果有人威胁对你使用这些武器,你总是可以威胁进行核反击。我担心的是,我不完全确定我们能否对数据中心里的天才国度的核威慑充满信心:强大的 AI 可能设计出探测和打击核潜艇的方法,对核武器基础设施的操作人员进行影响行动,或利用 AI 的网络能力对用于探测核发射的卫星发动网络攻击。33

First, some people might put their faith in the nuclear deterrent, particularly to counter the use of AI autonomous weapons for military conquest. If someone threatens to use these weapons against you, you can always threaten a nuclear response back. My worry is that I’m not totally sure we can be confident in the nuclear deterrent against a country of geniuses in a datacenter: it is possible that powerful AI could devise ways to detect and strike nuclear submarines, conduct influence operations against the operators of nuclear weapons infrastructure, or use AI’s cyber capabilities to launch a cyberattack against satellites used to detect nuclear launches.33

33 当然,这也是一个改进核威慑安全性的论据,使其更有可能在强大 AI 面前保持稳健,拥有核武器的民主国家应该这样做。但我们不知道强大的 AI 将能够做什么,或者哪些防御措施(如果有的话)能有效对抗它,因此我们不应假设这些措施必然能解决问题。

33 This is, of course, also an argument for improving the security of the nuclear deterrent to make it more likely to be robust against powerful AI, and nuclear-armed democracies should do this. But we don’t know what a powerful AI will be capable of or which defenses, if any, will work against it, so we should not assume that these measures will necessarily solve the problem.

或者,可能仅凭 AI 监控和 AI 宣传就能接管国家,而从未出现一个明确的时刻让人看清发生了什么,核反击也显得不合适。_也许_这些事情不可行,核威慑仍然有效,但赌注太高,不能冒险。34

Alternatively, it’s possible that taking over countries is feasible with only AI surveillance and AI propaganda, and never actually presents a clear moment where it’s obvious what is going on and where a nuclear response would be appropriate. _Maybe_ these things aren’t feasible and the nuclear deterrent will still be effective, but it seems too high stakes to take a risk.34

34 还有一种风险是,即使核威慑仍然有效,攻击国可能决定试探我们的底线——不清楚我们是否愿意用核武器来防御无人机群,即使无人机群有相当大的风险征服我们。无人机群可能是一种新事物,比核攻击轻但比常规攻击重。或者,对 AI 时代核威慑有效性的不同评估可能以破坏稳定的方式改变核冲突的博弈论。

34 There is also the risk that even if the nuclear deterrent remains effective, an attacking country might decide to call our bluff—it’s unclear whether we’d be willing to use nuclear weapons to defend against a drone swarm even if the drone swarm has a substantial risk of conquering us. Drone swarms might be a new thing that is less severe than nuclear attacks but more severe than conventional attacks. Alternatively, differing assessments of the effectiveness of the nuclear deterrent in the age of AI might alter the game theory of nuclear conflict in a destabilizing manner.

第二个可能的反对意见是,我们可以对这些独裁工具采取反制措施。我们可以用自己的无人机对抗无人机,网络防御将随着网络攻击而改进,可能有办法使人们免受宣传影响等等。我的回应是,这些防御只有借助同等强大的 AI 才可能实现。如果没有一个同样聪明且数量众多的数据中心里的天才国度作为反制力量,就无法在质量或数量上匹配无人机,网络防御也无法智胜网络攻击等等。因此,反制措施的问题归结为强大 AI 中的力量平衡问题。在这里,我担心强大 AI 的递归或自我强化特性(我在本文开头讨论过):每一代 AI 都可以用来设计和训练下一代 AI。这导致了失控优势的风险,即当前强大 AI 的领先者可能能够扩大其领先优势,并且难以追赶。我们需要确保不是威权国家首先进入这个循环。

A second possible objection is that there might be countermeasures we can take against these tools of autocracy. We can counter drones with our own drones, cyberdefense will improve along with cyberattack, there may be ways to immunize people against propaganda, etc. My response is that these defenses will only be possible with comparably powerful AI. If there isn’t some counterforce with a comparably smart and numerous country of geniuses in a datacenter, it won’t be possible to match the quality or quantity of drones, for cyberdefense to outsmart cyberoffense, etc. So the question of countermeasures reduces to the question of a balance of power in powerful AI. Here, I am concerned about the recursive or self-reinforcing property of powerful AI (which I discussed at the beginning of this essay): that each generation of AI can be used to design and train the next generation of AI. This leads to a risk of a runaway advantage, where the current leader in powerful AI may be able to increase their lead and may be difficult to catch up with. We need to make sure it is not an authoritarian country that gets to this loop first.

此外,即使可以实现力量平衡,世界仍有可能分裂成独裁势力范围,就像《一九八四》中那样。即使几个竞争大国各自拥有其强大的 AI 模型,并且没有一个能压倒其他,每个大国仍然可以内部镇压其本国人民,并且很难被推翻(因为人民没有强大的 AI 来保护自己)。因此,即使没有导致单一国家统治世界,防止 AI 赋能的独裁也很重要。

Furthermore, even if a balance of power can be achieved, there is still risk that the world could be split up into autocratic spheres, as in _Nineteen Eighty-Four_. Even if several competing powers each have their powerful AI models, and none can overpower the others, each power could still internally repress their own population, and would be very difficult to overthrow (since the populations don’t have powerful AI to defend themselves). It is thus important to prevent AI-enabled autocracy even if it doesn’t lead to a single country taking over the world.

防御措施 Defenses

我们如何防御这一系列广泛的专制工具和潜在威胁行为者?正如前几节所述,我认为有几件事可以做。首先,我们绝对不应该向中共出售芯片、芯片制造工具或数据中心。芯片和芯片制造工具是强大 AI 的最大瓶颈,阻止它们是一个简单但极其有效的措施,也许是我们能采取的最重要的单一行动。向中共出售用于构建 AI 极权国家并可能军事征服我们的工具是毫无意义的。有许多复杂的论点试图为这类销售辩护,例如“将我们的技术栈传播到世界各地”能让“美国在某种未明确的经济战斗中获胜”。在我看来,这就像向朝鲜出售核武器,然后吹嘘导弹外壳由波音制造,因此美国“赢了”。中国在批量生产前沿芯片的能力上落后美国数年,而在数据中心构建天才国家的关键时期很可能就在接下来的几年内。

How do we defend against this wide range of autocratic tools and potential threat actors? As in the previous sections, there are several things I think we can do. First, we should absolutely not be selling chips, chip-making tools, or datacenters to the CCP. Chips and chip-making tools are the single greatest bottleneck to powerful AI, and blocking them is a simple but extremely effective measure, perhaps the most important single action we can take. It makes no sense to sell the CCP the tools with which to build an AI totalitarian state and possibly conquer us militarily. A number of complicated arguments are made to justify such sales, such as the idea that “spreading our tech stack around the world” allows “America to win” in some general, unspecified economic battle. In my view, this is like selling nuclear weapons to North Korea and then bragging that the missile casings are made by Boeing and so the US is “winning.” China is several years behind the US in their ability to produce frontier chips in quantity, and the critical period for building the country of geniuses in a datacenter is very likely to be within those next several years.35

明确地说,即使强大 AI 的时间线大幅延长,我也认为不向中国出售芯片是正确的策略。我们无法让中国“上瘾”于美国芯片——他们决心以某种方式发展自己的本土芯片产业。这将花费他们许多年,而我们出售芯片所做的只是在此期间给他们一个巨大的推动。

35 To be clear, I would believe it is the right strategy not to sell chips to China, even if the timeline to powerful AI were substantially longer. We cannot get the Chinese “addicted” to American chips—they are determined to develop their native chip industry one way or another. It will take them many years to do so, and all we are doing by selling them chips is giving them a big boost during that time.

没有理由在这个关键时期给他们的 AI 产业一个巨大的推动。

There is no reason to give a giant boost to their AI industry during this critical period.

第二,利用 AI 赋能民主国家以抵抗专制国家是有意义的。这就是 Anthropic 认为向美国及其民主盟友的情报和国防社区提供 AI 很重要的原因。保护受到攻击的民主国家,如乌克兰和(通过网络攻击的)台湾,似乎是特别优先的事项,同样重要的是赋能民主国家利用其情报服务从内部瓦解和削弱专制国家。在某种程度上,应对专制威胁的唯一方法是在军事上匹敌并超越他们。美国及其民主盟友的联盟,如果在强大 AI 方面取得主导地位,将不仅能自卫对抗专制国家,还能遏制它们并限制其 AI 极权滥用。

Second, it makes sense to use AI to empower democracies to resist autocracies. This is the reason Anthropic considers it important to provide AI to the intelligence and defense communities in the US and its democratic allies. Defending democracies that are under attack, such as Ukraine and (via cyber attacks) Taiwan, seems especially high priority, as does empowering democracies to use their intelligence services to disrupt and degrade autocracies from the inside. At some level the only way to respond to autocratic threats is to match and outclass them militarily. A coalition of the US and its democratic allies, if it achieved predominance in powerful AI, would be in a position to not only defend itself against autocracies, but contain them and limit their AI totalitarian abuses.

第三,我们需要在民主国家内部对 AI 滥用划定严格界限。我们必须限制政府利用 AI 的行为,以防止他们夺取权力或压迫自己的人民。我提出的原则是:我们应以所有方式将 AI 用于国防,但那些会使我们更像专制对手的方式除外。

Third, we need to draw a hard line against AI abuses within democracies. There need to be limits to what we allow our governments to do with AI, so that they don’t seize power or repress their own people. The formulation I have come up with is that we should use AI for national defense in all ways _except those which would make us more like our autocratic adversaries_.

界限应划在哪里?在本节开头的列表中,有两个项目——将 AI 用于国内大规模监控和大规模宣传——在我看来是明确的红线,完全不可接受。有人可能认为没有必要采取行动(至少在美国),因为国内大规模监控根据第四修正案已经违法。但 AI 的快速进步可能创造出我们现有法律框架难以应对的情况。例如,美国政府大规模记录所有公共对话(例如人们在街角互相说的话)可能并不违宪,而以前很难处理如此大量的信息,但借助 AI,所有这些都可以被转录、解释和三角定位,从而描绘出许多或大多数公民的态度和忠诚度。我会支持以公民自由为重点的立法(甚至可能是宪法修正案),以对 AI 驱动的滥用行为设置更强的防护栏。

Where should the line be drawn? In the list at the beginning of this section, two items—using AI for domestic mass surveillance and mass propaganda—seem to me like bright red lines and entirely illegitimate. Some might argue that there’s no need to do anything (at least in the US), since domestic mass surveillance is already illegal under the Fourth Amendment. But the rapid progress of AI may create situations that our existing legal frameworks are not well designed to deal with. For example, it would likely not be unconstitutional for the US government to conduct massively scaled recordings of all _public_ conversations (e.g., things people say to each other on a street corner), and previously it would have been difficult to sort through this volume of information, but with AI it could all be transcribed, interpreted, and triangulated to create a picture of the attitude and loyalties of many or most citizens. I would support civil liberties-focused legislation (or maybe even a constitutional amendment) that imposes stronger guardrails against AI-powered abuses.

另外两项——完全自主武器和用于战略决策的 AI——更难划定界限,因为它们在捍卫民主方面有合法用途,同时也容易被滥用。我认为这里需要的是极度谨慎和审查,同时设置防护栏以防止滥用。我主要的担忧是“按按钮的手指”数量太少,以至于一个人或少数几个人可以在没有其他人合作执行命令的情况下操作一支无人机军队。随着 AI 系统变得更强大,我们可能需要更直接和即时的监督机制来确保它们不被滥用,可能涉及行政分支以外的政府分支。我认为我们应特别谨慎地对待完全自主武器,

The other two items—fully autonomous weapons and AI for strategic decision-making—are harder lines to draw since they have legitimate uses in defending democracy, while also being prone to abuse. Here I think what is warranted is extreme care and scrutiny combined with guardrails to prevent abuses. My main fear is having too small a number of “fingers on the button,” such that one or a handful of people could essentially operate a drone army without needing any other humans to cooperate to carry out their orders. As AI systems get more powerful, we may need to have more direct and immediate oversight mechanisms to ensure they are not misused, perhaps involving branches of government other than the executive. I think we should approach fully autonomous weapons in particular with great caution,36

明确地说,目前在乌克兰和台湾使用的大多数武器并非完全自主武器。这些即将到来,但尚未出现。

36 To be clear, most of what is being used in Ukraine and Taiwan today are not _fully_ autonomous weapons. These are coming, but not here today.

并且在没有适当保障措施的情况下不要急于使用它们。

and not rush into their use without proper safeguards.

第四,在民主国家内部对 AI 滥用划定严格界限后,我们应利用这一先例,针对强大 AI 的最严重滥用行为建立国际禁忌。我认识到当前政治风向已转向反对国际合作和国际规范,但在这个案例中我们迫切需要它们。世界需要了解强大 AI 在专制者手中的黑暗潜力,并认识到某些 AI 使用等同于试图永久窃取他们的自由并强加一个他们无法逃脱的极权国家。我甚至认为,在某些情况下,利用强大 AI 进行大规模监控、大规模宣传以及某些类型的进攻性完全自主武器使用应被视为反人类罪。更一般地说,迫切需要建立反对 AI 赋能极权主义及其所有工具和手段的强有力规范。

Fourth, after drawing a hard line against AI abuses in democracies, we should use that precedent to create an international taboo against the worst abuses of powerful AI. I recognize that the current political winds have turned against international cooperation and international norms, but this is a case where we sorely need them. The world needs to understand the dark potential of powerful AI in the hands of autocrats, and to recognize that certain uses of AI amount to an attempt to permanently steal their freedom and impose a totalitarian state from which they can’t escape. I would even argue that in some cases, large-scale surveillance with powerful AI, mass propaganda with powerful AI, and certain types of _offensive_ uses of fully autonomous weapons should be considered crimes against humanity. More generally, a robust norm against AI-enabled totalitarianism and all its tools and instruments is sorely needed.

甚至可以有更强版本的立场:由于 AI 赋能极权主义的可能性如此黑暗,专制主义根本就不是人们在强大 AI 时代后可以接受的政府形式。正如封建主义在工业革命中变得不可行一样,AI 时代可能不可避免地导致这样的结论:如果人类要拥有美好的未来,民主(并且希望是经过 AI 改进和振兴的民主,正如我在《仁慈的机器》中所讨论的)是唯一可行的政府形式。

It is possible to have an even stronger version of this position, which is that because the possibilities of AI-enabled totalitarianism are so dark, autocracy is simply not a form of government that people can accept in the post-powerful AI age. Just as feudalism became unworkable with the industrial revolution, the AI age could lead inevitably and logically to the conclusion that democracy (and, hopefully, democracy improved and reinvigorated by AI, as I discuss in _Machines of Loving Grace_) is the only viable form of government if humanity is to have a good future.

第五,也是最后一点,AI 公司应受到密切关注,它们与政府的联系也是必要的,但必须有界限和限制。强大 AI 所蕴含的巨大能力使得普通的公司治理——旨在保护股东并防止欺诈等普通滥用行为——不太可能胜任治理 AI 公司的任务。公司公开承诺(甚至作为公司治理的一部分)不采取某些行动也可能有价值,例如私下建造或囤积军事硬件,由单个人以不负责任的方式使用大量计算资源,或利用其 AI 产品作为宣传工具来操纵公众舆论以利于自己。

Fifth and finally, AI companies should be carefully watched, as should their connection to the government, which is necessary, but must have limits and boundaries. The sheer amount of capability embodied in powerful AI is such that ordinary corporate governance—which is designed to protect shareholders and prevent ordinary abuses such as fraud—is unlikely to be up to the task of governing AI companies. There may also be value in companies publicly committing to (perhaps even as part of corporate governance) not take certain actions, such as privately building or stockpiling military hardware, using large amounts of computing resources by single individuals in unaccountable ways, or using their AI products as propaganda to manipulate public opinion in their favor.

危险来自许多方向,有些方向相互矛盾。唯一不变的是,我们必须为每个人寻求问责、规范和防护栏,即使我们赋能“好”行为者以制约“坏”行为者。

The danger here comes from many directions, and some directions are in tension with others. The only constant is that we must seek accountability, norms, and guardrails for everyone, even as we empower “good” actors to keep “bad” actors in check.

经济冲击 Economic disruption

前三节主要讨论了强大 AI 带来的安全风险:来自 AI 本身的风险、来自个人和小组织滥用的风险,以及来自国家和大型组织滥用的风险。如果我们暂且搁置安全风险,或者假设它们已被解决,那么下一个问题就是经济问题。这种惊人的“人力”资本的注入将对经济产生什么影响?显然,最直接的影响将是极大地促进经济增长。科学研究、生物医学创新、制造业、供应链、金融体系效率等领域的进步速度,几乎必然会导致经济增长率大幅提升。在《爱之机器》一书中,我提出可能实现 10%–20%的持续年 GDP 增长率。

The previous three sections were essentially about security risks posed by powerful AI: risks from the AI itself, risks from misuse by individuals and small organizations and risks of misuse by states and large organizations. If we put aside security risks or assume they have been solved, the next question is economic. What will be the effect of this infusion of incredible “human” capital on the economy? Clearly, the most obvious effect will be to greatly increase economic growth. The pace of advances in scientific research, biomedical innovation, manufacturing, supply chains, the efficiency of the financial system, and much more are almost guaranteed to lead to a much faster rate of economic growth. In _Machines of Loving Grace_, I suggest that a 10–20% sustained annual GDP growth rate may be possible.

但显而易见,这是一把双刃剑:在这样的世界中,大多数现有人类的经济前景如何?新技术常常带来劳动力市场冲击,而过去人类总能从中恢复,但我担心这是因为以往冲击只影响了人类全部能力范围中的一小部分,为人类拓展到新任务留下了空间。AI 的影响将更为广泛且发生得更快,因此我担心要让一切顺利运转将更具挑战性。

But it should be clear that this is a double-edged sword: what are the economic prospects for most existing humans in such a world? New technologies often bring labor market shocks, and in the past humans have always recovered from them, but I am concerned that this is because these previous shocks affected only a small fraction of the full possible range of human abilities, leaving room for humans to expand to new tasks. AI will have effects that are much broader and occur much faster, and therefore I worry it will be much more challenging to make things work out well.

劳动力市场冲击 Labor market disruption

我担心两个具体问题:劳动力市场替代和经济权力集中。先谈第一个。这是我在 2025 年非常公开地警告过的话题,当时我预测,即使 AI 加速经济增长和科学进步,它也可能在未来 1-5 年内取代一半的初级白领工作。这一警告引发了关于该话题的公开辩论。许多 CEO、技术专家和经济学家同意我的观点,但其他人认为我陷入了“劳动总量固定”的谬误,不了解劳动力市场如何运作,还有一些人没有注意到 1-5 年的时间范围,以为我在声称 AI 正在立即取代工作(我同意这很可能不是事实)。因此,有必要详细说明我为何担心劳动力替代,以澄清这些误解。

There are two specific problems I am worried about: labor market displacement, and concentration of economic power. Let’s start with the first one. This is a topic that I warned about very publicly in 2025, where I predicted that AI could displace half of all entry-level white collar jobs in the next 1–5 years, even as it accelerates economic growth and scientific progress. This warning started a public debate about the topic. Many CEOs, technologists, and economists agreed with me, but others assumed I was falling prey to a “lump of labor” fallacy and didn’t know how labor markets worked, and some didn’t see the 1–5-year time range and thought I was claiming AI is displacing jobs right now (which I agree it is likely not). So it is worth going through in detail why I am worried about labor displacement, to clear up these misunderstandings.

作为基线,了解劳动力市场通常如何应对技术进步是有用的。当新技术出现时,它首先使人类工作的某些部分变得更高效。例如,在工业革命早期,机器(如改进的犁)使农民在某些工作方面更高效。这提高了农民的生产力,从而增加了他们的工资。

As a baseline, it’s useful to understand how labor markets _normally_ respond to advances in technology. When a new technology comes along, it starts by making pieces of a given human job more efficient. For example, early in the Industrial Revolution, machines, such as upgraded plows, enabled human farmers to be more efficient at some aspects of the job. This improved the productivity of farmers, which increased their wages.

下一步,工作的某些部分可以完全由机器完成,例如通过发明脱粒机或播种机。在这个阶段,人类完成的工作比例越来越低,但他们完成的工作因与机器互补而变得更具杠杆效应,生产力持续上升。正如杰文斯悖论所述,农民的工资甚至农民的数量可能继续增加。即使 90%的工作由机器完成,人类只需将剩余 10%的工作量增加 10 倍,就能用同样的劳动产出 10 倍的产量。

In the next step, some parts of the job of farming could be done _entirely_ by machines, for example with the invention of the threshing machine or seed drill. In this phase, humans did a lower and lower fraction of the job, but the work they _did_ complete became more and more leveraged because it is complementary to the work of machines, and their productivity continued to rise. As described by Jevons’ paradox, the wages of farmers and perhaps even the number of farmers continued to increase. Even when 90% of the job is being done by machines, humans can simply do 10x more of the 10% they still do, producing 10x as much output for the same amount of labor.

最终,机器完成全部或几乎所有工作,就像现代联合收割机、拖拉机和其他设备一样。此时,作为人类就业形式的农业确实急剧下降,短期内可能造成严重冲击,但由于农业只是人类能够从事的众多有用活动之一,人们最终会转向其他工作,如操作工厂机器。即使农业事前占就业的很大比例,情况也是如此。250 年前,90%的美国人生活在农场;在欧洲,50-60%的就业是农业。现在这些地方的百分比降至个位数,因为工人转向了工业工作(后来是知识工作)。经济可以用仅 1-2%的劳动力完成以前需要大部分劳动力才能完成的工作,从而解放其余劳动力去建设更先进的工业社会。不存在固定的“劳动总量”,只有用更少资源做更多事情的不断扩展的能力。一旦短期冲击过去,人们的工资随 GDP 指数增长而上升,经济保持充分就业。

Eventually, machines do everything or almost everything, as with modern combine harvesters, tractors, and other equipment. At this point farming as a form of human employment really does go into steep decline, and this potentially causes serious disruption in the short term, but because farming is just one of many useful activities that humans are able to do, people eventually switch to other jobs, such as operating factory machines. This is true even though farming accounted for a huge proportion of employment _ex ante_. 250 years ago, 90% of Americans lived on farms; in Europe, 50–60% of employment was agricultural. Now those percentages are in the low single digits in those places, because workers switched to industrial jobs (and later, knowledge work jobs). The economy can do what previously required most of the labor force with only 1–2% of it, freeing up the rest of the labor force to build an ever more advanced industrial society. There’s no fixed “lump of labor,” just an ever-expanding ability to do more and more with less and less. People’s wages rise in line with the GDP exponential and the economy maintains full employment once disruptions in the short term have passed.

AI 的情况可能大致相同,但我强烈反对这种可能性。以下是我认为 AI 可能不同的几个原因:

It’s possible things will go roughly the same way with AI, but I would bet pretty strongly against it. Here are some reasons I think AI is likely to be different:

* **速度。** AI 进步的速度远快于以往的技术革命。例如,在过去两年中,AI 模型从几乎无法完成一行代码,发展到为某些人(包括 Anthropic 的工程师)编写全部或几乎所有代码。3737 我们最新模型 Claude Opus 4.5 的模型卡显示,Opus 在 Anthropic 经常使用的性能工程面试中表现优于公司历史上任何面试者。很快,它们可能端到端地完成软件工程师的全部任务。3838 “编写所有代码”和“端到端完成软件工程师的任务”是非常不同的,因为软件工程师做的远不止编写代码,还包括测试、处理环境、文件和安装、管理云算力部署、迭代产品等。人们很难适应这种变化速度,无论是适应特定工作方式的变化,还是转向新工作的需求。即使是传奇程序员也越来越多地称自己“落后”。随着 AI 编码模型加速 AI 开发任务,速度可能还会继续加快。需要明确的是,速度本身并不意味着劳动力市场和就业最终不会恢复,它只是意味着与过去的技术相比,短期转型将异常痛苦,因为人类和劳动力市场反应和均衡的速度很慢。

* Speed.The pace of progress in AI is much faster than for previous technological revolutions. For example, in the last 2 years, AI models went from barely being able to complete a single line of code, to writing all or almost all of the code for some people—including engineers at Anthropic.3737 Our model card for Claude Opus 4.5, our most recent model, shows that Opus performs better on a performance engineering interview frequently given at Anthropic than any interviewee in the history of the company. Soon, they may do the entire task of a software engineer end to end.3838 “Writing all of the code” and “doing the task of a software engineer end to end” are very different things, because software engineers do much more than just write code, including testing, dealing with environments, files, and installation, managing cloud compute deployments, iterating on products, and much more. It is hard for people to adapt to this pace of change, both to the changes in how a given job works and in the need to switch to new jobs. Even legendary programmers are increasingly describing themselves as “behind.” The pace may if anything continue to speed up, as AI coding models increasingly accelerate the task of AI development. To be clear, speed in itself does not mean labor markets and employment won’t eventually recover, it just implies the short-term transition will be unusually painful compared to past technologies, since humans and labor markets are slow to react and to equilibrate.

* **认知广度。** 正如“数据中心里的天才国度”这一说法所暗示的,AI 将能够完成非常广泛的人类认知能力——也许是全部。这与机械化农业、运输甚至计算机等以往技术非常不同。3939 计算机在某种意义上是通用的,但显然无法独立完成绝大多数人类认知能力,尽管在少数领域(如算术)远超人类。当然,建立在计算机之上的东西(如 AI)现在能够完成广泛的认知能力,这正是本文的主题。这将使人们更难从被替代的工作轻松转向他们适合的类似工作。例如,金融、咨询和法律等领域的初级工作所需的一般智力能力相当相似,即使具体知识差异很大。一项只颠覆其中一种的技术会让员工转向其他两种近似替代(或让本科生换专业)。但同时颠覆所有三种(以及许多其他类似工作)可能更难适应。此外,不仅仅是大多数现有工作将被颠覆。这种情况以前发生过——回想一下农业曾占就业的很大比例。但农民可以转向相对相似的工厂机器操作工作,即使这种工作以前并不常见。相比之下,AI 越来越匹配人类的一般认知特征,这意味着它也将擅长那些通常为应对旧工作被自动化而创造的新工作。另一种说法是,AI 不是特定人类工作的替代品,而是人类的通用劳动替代品。

* Cognitive breadth.As suggested by the phrase “country of geniuses in a datacenter,” AI will be capable of a very wide range of human cognitive abilities—perhaps all of them. This is very different from previous technologies like mechanized farming, transportation, or even computers.3939 Computers are general in a sense, but are clearly incapable on their own of the vast majority of human cognitive abilities, even as they greatly exceed humans in a few areas (such as arithmetic). Of course, things built _on top_ of computers, such as AI, are now capable of a wide range of cognitive abilities, which is what this essay is about. This will make it harder for people to switch easily from jobs that are displaced to similar jobs that they would be a good fit for. For example, the general intellectual abilities required for entry-level jobs in, say, finance, consulting, and law are fairly similar, even if the specific knowledge is quite different. A technology that disrupted only one of these three would allow employees to switch to the two other close substitutes (or for undergraduates to switch majors). But disrupting all three at once (along with many other similar jobs) may be harder for people to adapt to. Furthermore, it’s not _just_ that most existing jobs will be disrupted. That part has happened before—recall that farming was a huge percentage of employment. But farmers could switch to the relatively similar work of operating factory machines, even though that work hadn’t been common before. By contrast, AI is increasingly matching the general cognitive profile of humans, which means it will also be good at the new jobs that would ordinarily be created in response to the old ones being automated. Another way to say it is that AI isn’t a substitute for specific human jobs but rather a general labor substitute for humans.

* **按认知能力分层。** 在广泛的任务中,AI 似乎从能力阶梯的底部向顶部推进。例如,在编程方面,我们的模型从“平庸的程序员”水平发展到“强大的程序员”,再到“非常强大的程序员”。4040 需要明确的是,AI 模型与人类的能力强弱分布并不完全相同。但它们也在每个维度上相当均匀地进步,因此尖峰或不均匀的分布可能最终无关紧要。我们现在开始在白领工作中普遍看到同样的进展。因此,我们面临一种风险:AI 不是影响具有特定技能或特定职业的人(他们可以通过再培训适应),而是影响具有某些内在认知属性(即较低智力能力,更难改变)的人。不清楚这些人将去哪里或做什么,我担心他们可能形成一个失业或极低工资的“底层阶级”。需要明确的是,类似情况以前发生过——例如,一些经济学家认为计算机和互联网代表了“技能偏向型技术变革”。但这种技能偏向既没有我预期 AI 带来的那么极端,也被认为加剧了工资不平等,4141 尽管经济学家对此存在争议。 所以这并非令人安心的先例。

* Slicing by cognitive ability.Across a wide range of tasks, AI appears to be advancing from the bottom of the ability ladder to the top. For example, in coding our models have proceeded from the level of “a mediocre coder” to “a strong coder” to “a very strong coder.”4040 To be clear, AI models do not have precisely the same profile of strengths and weaknesses as humans. But they are also advancing fairly uniformly along every dimension, such that having a spiky or uneven profile may not ultimately matter. We are now starting to see the same progression in white-collar work in general. We are thus at risk of a situation where, instead of affecting people with specific skills or in specific professions (who can adapt by retraining), AI is affecting people with certain intrinsic cognitive properties, namely lower intellectual ability (which is harder to change). It is not clear where these people will go or what they will do, and I am concerned that they could form an unemployed or very-low-wage “underclass.” To be clear, things somewhat like this have happened before—for example, computers and the internet are believed by some economists to represent “skill-biased technological change.” But this skill biasing was both not as extreme as what I expect to see with AI, and is believed to have contributed to an increase in wage inequality,4141 Though there is debateamongeconomists about this idea. so it is not exactly a reassuring precedent.

* **填补空白的能力。** 人类工作面对新技术时的调整方式通常是,工作有许多方面,而新技术即使看似直接替代人类,也往往存在空白。如果有人发明了制造小部件的机器,人类可能仍需将原材料装入机器。即使这只需要手工制造小部件 1%的努力,人类工人也可以简单地将产量提高 100 倍。但 AI 除了是快速进步的技术外,也是快速适应的技术。在每次模型发布期间,AI 公司都会仔细衡量模型的优劣,客户在发布后也会提供此类信息。弱点可以通过收集体现当前空白的任务并在下一个模型上训练来解决。在生成式 AI 早期,用户注意到 AI 系统存在某些弱点(如 AI 图像模型生成的手指数量错误),许多人认为这些弱点是技术固有的。如果是这样,就会限制工作替代。但几乎每个这样的弱点都会很快得到解决——通常只需几个月。

* Ability to fill in the gaps.The way human jobs often adjust in the face of new technology is that there are many aspects to the job, and the new technology, even if it appears to directly replace humans, often has gaps in it. If someone invents a machine to make widgets, humans may still have to load raw material into the machine. Even if that takes only 1% as much effort as making the widgets manually, human workers can simply make 100x more widgets. But AI, in addition to being a rapidly advancing technology, is also a rapidly _adapting_ technology. During every model release, AI companies carefully measure what the model is good at and what it isn’t, and customers also provide such information after the launch. Weaknesses can be addressed by collecting tasks that embody the current gap, and training on them for the next model. Early in generative AI, users noticed that AI systems had certain weaknesses (such as AI image models generating hands with the wrong number of fingers) and many assumed these weaknesses were inherent to the technology. If they were, it would limit job disruption. But pretty much every such weakness gets addressed quickly— often, within just a few months.

值得讨论常见的质疑点。首先,有人认为经济扩散会很慢,因此即使底层技术能够完成大部分人类劳动,它在整个经济中的实际应用可能慢得多(例如在远离 AI 行业且采用缓慢的行业)。技术扩散缓慢确实是事实——我与来自各种企业的人交谈,有些地方采用 AI 需要数年时间。这就是为什么我预测 50%的初级白领工作被替代需要 1-5 年,尽管我怀疑我们将在不到 5 年内拥有强大的 AI(从技术上讲足以完成大部分或所有工作,而不仅仅是初级工作)。但扩散效应只是为我们争取时间。而且我不确定它们会像人们预测的那样慢。企业 AI 采用率远高于以往任何技术,主要得益于技术本身的强大力量。此外,即使传统企业采用新技术缓慢,初创公司也会涌现出来充当“粘合剂”,使采用更容易。如果这不起作用,初创公司可能直接颠覆现有企业。

It’s worth addressing common points of skepticism. First, there is the argument that economic diffusion will be slow, such that even if the underlying technology is _capable_ of doing most human labor, the actual application of it across the economy may be much slower (for example in industries that are far from the AI industry and slow to adopt). Slow diffusion of technology is definitely real—I talk to people from a wide variety of enterprises, and there are places where the adoption of AI will take years. That’s why my prediction for 50% of entry level white collar jobs being disrupted is 1–5 years, even though I suspect we’ll have powerful AI (which would be, technologically speaking, enough to do _most or all_ jobs, not just entry level) in much less than 5 years. But diffusion effects merely buy us time. And I am not confident they will be as slow as people predict. Enterprise AI adoption is growing at rates much faster than any previous technology, largely on the pure strength of the technology itself. Also, even if traditional enterprises are slow to adopt new technology, startups will spring up to serve as “glue” and make the adoption easier. If that doesn’t work, the startups may simply disrupt the incumbents directly.

这可能导致一个世界,与其说是特定工作被替代,不如说是大型企业普遍被颠覆,并被劳动力密集度低得多的初创公司取代。这也可能导致一个“地理不平等”的世界,越来越多的世界财富集中在硅谷,硅谷成为自己的经济体,以不同于世界其他地区的速度运行,并将其抛在后面。所有这些结果对经济增长都有好处——但对劳动力市场或被抛在后面的人则不那么好。

That could lead to a world where it isn’t so much that specific jobs are disrupted as it is that large enterprises are disrupted in general and replaced with much less labor-intensive startups. This could also lead to a world of “geographic inequality,” where an increasing fraction of the world’s wealth is concentrated in Silicon Valley, which becomes its own economy running at a different speed than the rest of the world and leaving it behind. All of these outcomes would be great for economic growth—but not so great for the labor market or those who are left behind.

其次,有些人说人类工作将转向物理世界,从而避开 AI 快速进步的“认知劳动”范畴。我也不确定这有多安全。许多体力劳动已经由机器完成(如制造业),或很快将由机器完成(如驾驶)。此外,足够强大的 AI 将能够加速机器人开发,然后控制这些机器人在物理世界中工作。这可能争取一些时间(这是好事),但我担心争取不了太多。即使冲击仅限于认知任务,它仍然是前所未有的巨大而迅速的冲击。

Second, some people say that human jobs will move to the physical world, which avoids the whole category of “cognitive labor” where AI is progressing so rapidly. I am not sure how safe this is, either. A lot of physical labor is already being done by machines (e.g., manufacturing) or will soon be done by machines (e.g., driving). Also, sufficiently powerful AI will be able to accelerate the development of robots, and then control those robots in the physical world. It may buy some time (which is a good thing), but I’m worried it won’t buy much. And even if the disruption was limited only to cognitive tasks, it would still be an unprecedentedly large and rapid disruption.

第三,也许有些任务本质上需要或大大受益于人类接触。我对这一点更不确定,但我仍然怀疑它是否足以抵消我上面描述的大部分影响。AI 已广泛用于客户服务。许多人报告说,与 AI 谈论个人问题比与治疗师交谈更容易——AI 更有耐心。当我姐姐在怀孕期间遇到医疗问题时,她觉得没有从护理人员那里得到所需的答案或支持,她发现 Claude 有更好的床边态度(并且在诊断问题上也更成功)。我确信有些任务人类接触确实重要,但我不确定有多少——而且我们在这里谈论的是为劳动力市场中几乎每个人找到工作。

Third, perhaps some tasks inherently require or greatly benefit from a human touch. I’m a little more uncertain about this one, but I’m still skeptical that it will be enough to offset the bulk of the impacts I described above. AI is already widely used for customer service. Many people report that it is easier to talk to AI about their personal problems than to talk to a therapist—that the AI is more patient. When my sister was struggling with medical problems during a pregnancy, she felt she wasn’t getting the answers or support she needed from her care providers, and she found Claude to have a better bedside manner (as well as succeeding better at diagnosing the problem). I’m sure there are some tasks for which a human touch really is important, but I’m not sure how many—and here we’re talking about finding work for nearly everyone in the labor market.

第四,有些人可能认为比较优势仍将保护人类。根据比较优势定律,即使 AI 在所有方面都比人类强,人类与 AI 技能分布的任何相对差异都会在人类和 AI 之间创造贸易和专业化基础。问题是,如果 AI 的生产力比人类高出数千倍,这种逻辑就开始失效。即使是微小的交易成本也可能使 AI 与人类交易变得不值得。而人类的工资可能非常低,即使他们技术上仍有东西可提供。

Fourth, some may argue that comparative advantage will still protect humans. Under the law of comparative advantage, even if AI is better than humans at everything, any _relative_ differences between the human and AI profile of skills creates a basis of trade and specialization between humans and AI. The problem is that if AIs are literally thousands of times more productive than humans, this logic starts to break down. Even tiny transaction costs could make it not worth it for AI to trade with humans. And human wages may be very low, even if they technically have something to offer.

所有这些因素都有可能被解决——劳动力市场足够有韧性,能够适应如此巨大的冲击。但即使最终能适应,上述因素也表明短期冲击的规模将是前所未有的。

It’s possible all of these factors can be addressed—that the labor market is resilient enough to adapt to even such an enormous disruption. But even if it can eventually adapt, the factors above suggest that the short-term shock will be unprecedented in size.

防御措施 Defenses

我们该如何应对这个问题?我有几点建议,其中一些 Anthropic 已经在实施。首先,要实时获取关于就业岗位流失的准确数据。当经济变化非常迅速时,很难获得可靠的数据,而没有可靠的数据,就很难设计有效的政策。例如,政府数据目前缺乏关于各企业和行业采用 AI 的细粒度、高频数据。过去一年,Anthropic 一直在运营并公开发布一个经济指数,几乎实时显示我们模型的使用情况,并按行业、任务、地点甚至任务是被自动化还是协作完成等维度进行细分。我们还设立了一个经济咨询委员会,帮助我们解读这些数据并预测未来趋势。

What can we do about this problem? I have several suggestions, some of which Anthropic is already doing. The first thing is simply to get accurate data about what is happening with job displacement in real time. When an economic change happens very quickly, it’s hard to get reliable data about what is happening, and without reliable data it is hard to design effective policies. For example, government data is currently lacking granular, high-frequency data on AI adoption across firms and industries. For the last year Anthropic has been operating and publicly releasing an Economic Index that shows use of our models almost in real time, broken down by industry, task, location, and even things like whether a task was being automated or conducted collaboratively. We also have an Economic Advisory Council to help us interpret this data and see what is coming.

其次,AI 公司在与企业合作的方式上有选择余地。传统企业的低效率意味着它们部署 AI 的过程可能非常依赖路径,因此存在选择更好路径的空间。企业通常面临“成本节约”(用更少的人做同样的事)和“创新”(用同样的人做更多的事)之间的选择。市场最终会同时产生这两种结果,任何有竞争力的 AI 公司都必须兼顾两者,但或许存在一定的空间,可以在可能的情况下引导企业走向创新,这或许能为我们争取一些时间。Anthropic 正在积极思考这一点。

Second, AI companies have a choice in how they work with enterprises. The very inefficiency of traditional enterprises means that their rollout of AI can be very path dependent, and there is some room to choose a better path. Enterprises often have a choice between “cost savings” (doing the same thing with fewer people) and “innovation” (doing more with the same number of people). The market will inevitably produce both eventually, and any competitive AI company will have to serve some of both, but there may be some room to steer companies towards innovation when possible, and it may buy us some time. Anthropic is actively thinking about this.

第三,公司应该考虑如何照顾员工。短期内,在公司内部创造性地重新分配员工,可能是避免裁员的一种有希望的方式。长期来看,在一个总财富巨大、许多公司因生产率和资本集中而价值大幅提升的世界里,即使在员工不再提供传统意义上的经济价值之后,继续支付他们工资也可能是可行的。Anthropic 目前正在考虑一系列可能的方案,并将在近期分享。

Third, companies should think about how to take care of their employees. In the short term, being creative about ways to reassign employees within companies may be a promising way to stave off the need for layoffs. In the long term, in a world with enormous total wealth, in which many companies increase greatly in value due to increased productivity and capital concentration, it may be feasible to pay human employees even long after they are no longer providing economic value in the traditional sense. Anthropic is currently considering a range of possible pathways for our own employees that we will share in the near future.

第四,富人有义务帮助解决这个问题。令我感到悲哀的是,许多富人(尤其是科技行业的)最近采取了一种愤世嫉俗和虚无主义的态度,认为慈善事业不可避免地是欺诈或无用的。无论是像盖茨基金会这样的私人慈善机构,还是像 PEPFAR 这样的公共项目,都在发展中国家挽救了数千万人的生命,并在发达国家创造了经济机会。Anthropic 的所有联合创始人都承诺捐出我们 80%的财富,Anthropic 的员工个人也承诺捐出价值数十亿美元的公司股份——公司已承诺对这些捐款进行匹配。

Fourth, wealthy individuals have an obligation to help solve this problem. It is sad to me that many wealthy individuals (especially in the tech industry) have recently adopted a cynical and nihilistic attitude that philanthropy is inevitably fraudulent or useless. Both private philanthropy like the Gates Foundation and public programs like PEPFAR have saved tens of millions of lives in the developing world, and helped to create economic opportunity in the developed world. All of Anthropic’s co-founders have pledged to donate 80% of our wealth, and Anthropic’s staff have individually pledged to donate company shares worth billions at current prices—donations that the company has committed to matching.

第五,尽管上述所有私人行动都有帮助,但最终如此巨大的宏观经济问题需要政府干预。面对巨大的经济蛋糕与高度不平等(由于许多人缺乏工作或工作报酬低)并存的情况,自然的政策回应是累进税制。税收可以是普遍的,也可以专门针对 AI 公司。显然,税收设计很复杂,有很多出错的方式。我不支持设计糟糕的税收政策。我认为本文预测的极端不平等水平,从基本道德理由出发,证明更健全的税收政策是合理的,但我也可以向世界上的亿万富翁们提出一个务实的论点:支持一个好的版本符合他们的利益——如果他们不支持好的版本,他们最终会得到一个由暴民设计的糟糕版本。

Fifth, while all the above private actions can be helpful, ultimately a macroeconomic problem this large will require government intervention. The natural policy response to an enormous economic pie coupled with high inequality (due to a lack of jobs, or poorly paid jobs, for many) is progressive taxation. The tax could be general or could be targeted against AI companies in particular. Obviously tax design is complicated, and there are many ways for it to go wrong. I don’t support poorly designed tax policies. I think the extreme levels of inequality predicted in this essay justify a more robust tax policy on basic moral grounds, but I can also make a pragmatic argument to the world’s billionaires that it’s in their interest to support a good version of it: if they don’t support a good version, they’ll inevitably get a bad version designed by a mob.

最终,我认为上述所有干预措施都是争取时间的方式。最终,AI 将能够做所有事情,我们需要正视这一点。我希望到那时,我们可以利用 AI 本身来帮助重组市场,使其对每个人都有效,而上述干预措施可以帮助我们度过过渡期。

Ultimately, I think of all of the above interventions as ways to buy time. In the end AI will be able to do everything, and we need to grapple with that. It’s my hope that by that time, we can use AI itself to help us restructure markets in ways that work for everyone, and that the interventions above can get us through the transitional period.

经济权力集中 Economic concentration of power

与就业岗位流失或经济不平等本身不同,经济权力集中是另一个问题。第 1 节讨论了人类被 AI 剥夺权力的风险,第 3 节讨论了公民被政府通过武力或胁迫剥夺权力的风险。但另一种剥夺权力的情况可能发生:如果财富高度集中,一小群人凭借其影响力有效控制政府政策,而普通公民因缺乏经济杠杆而毫无影响力。民主最终依赖于这样一个理念:全体人口对经济运转是必要的。如果这种经济杠杆消失,民主的隐性社会契约可能不再有效。其他人已经对此有所论述,因此我无需在此详述,但我认同这种担忧,并担心它已经开始发生。

Separate from the problem of job displacement or economic inequality _per se_ is the problem of _economic concentration of power._ Section 1 discussed the risk that humanity gets disempowered by AI, and Section 3 discussed the risk that citizens get disempowered by their governments by force or coercion. But another kind of disempowerment can occur if there is such a huge concentration of wealth that a small group of people effectively controls government policy with their influence, and ordinary citizens have no influence because they lack economic leverage. Democracy is ultimately backstopped by the idea that the population as a whole is necessary for the operation of the economy. If that economic leverage goes away, then the implicit social contract of democracy may stop working. Others have written about this, so I needn’t go into great detail about it here, but I agree with the concern, and I worry it is already starting to happen.

需要明确的是,我并不反对人们赚很多钱。有强有力的论据表明,在正常情况下,这能激励经济增长。我理解有人担心扼杀创新的“金鹅”会阻碍创新。但在 GDP 年增长 10-20%、AI 迅速接管经济、而个人却持有 GDP 相当大份额的情况下,创新并不是需要担心的问题。需要担心的是财富集中到足以摧毁社会的程度。

To be clear, I am not opposed to people making a lot of money. There’s a strong argument that it incentivizes economic growth under normal conditions. I am sympathetic to concerns about impeding innovation by killing the golden goose that generates it. But in a scenario where GDP growth is 10–20% a year and AI is rapidly taking over the economy, yet single individuals hold appreciable fractions of the GDP, innovation is _not_ the thing to worry about. The thing to worry about is a level of wealth concentration that will break society.

美国历史上最著名的极端财富集中例子是镀金时代,而镀金时代最富有的实业家是约翰·D·洛克菲勒。洛克菲勒的财富约占当时美国 GDP 的 2%。42

The most famous example of extreme concentration of wealth in US history is the Gilded Age, and the wealthiest industrialist of the Gilded Age was John D. Rockefeller. Rockefeller’s wealth amounted to ~2% of the US GDP at the time.42

42 个人财富是“存量”,而 GDP 是“流量”,因此这并非声称洛克菲勒拥有美国经济价值的 2%。但衡量国家总财富比衡量 GDP 更难,且个人收入每年变化很大,因此很难用相同单位计算比率。最大个人财富与 GDP 的比率,虽然并非同类比较,但作为极端财富集中的基准仍然非常合理。

42 Personal wealth is a “stock,” while GDP is a “flow,” so this isn’t a claim that Rockefeller owned 2% of the economic value in the United States. But it’s harder to measure the total wealth of a nation than the GDP, and people’s individual incomes vary a lot per year, so it’s hard to make a ratio in the same units. The ratio of the largest personal fortune to GDP, while not comparing apples to apples, is nevertheless a perfectly reasonable benchmark for extreme wealth concentration.

类似的比例在今天将导致 6000 亿美元的财富,而当今世界首富(埃隆·马斯克)已经超过这一数字,约为 7000 亿美元。因此,即使在 AI 对经济产生大部分影响之前,我们已经处于历史上前所未有的财富集中水平。我认为(如果我们进入一个“天才之国”),不难想象 AI 公司、半导体公司以及可能的下游应用公司每年产生约 3 万亿美元的收入,43

A similar fraction today would lead to a fortune of $600B, and the richest person in the world today (Elon Musk) already exceeds that, at roughly $700B. So we are already at historically unprecedented levels of wealth concentration, even _before_ most of the economic impact of AI. I don’t think it is too much of a stretch (if we get a “country of geniuses”) to imagine AI companies, semiconductor companies, and perhaps downstream application companies generating ~$3T in revenue per year,43

43 整个经济中的劳动总价值为每年 60 万亿美元,因此每年 3 万亿美元相当于其中的 5%。如果一家公司以人类成本 20%的价格提供劳动力,并拥有 25%的市场份额,即使劳动力需求没有扩大(由于成本降低,几乎肯定会扩大),也能赚取这一数额。

43 The total value of labor across the economy is $60T/year, so $3T/year would correspond to 5% of this. That amount could be earned by a company that supplied labor for 20% of the cost of humans and had 25% market share, even if the demand for labor did not expand (which it almost certainly would due to the lower cost).

估值约 30 万亿美元,并导致个人财富达到数万亿美元。在那个世界里,我们今天关于税收政策的辩论将完全不适用,因为我们将处于一个根本不同的局面。

being valued at ~$30T, and leading to personal fortunes well into the trillions. In that world, the debates we have about tax policy today simply won’t apply as we will be in a fundamentally different situation.

与此相关的是,这种经济权力集中与政治体系的结合已经让我担忧。AI 数据中心已经占美国经济增长的相当大一部分,44

Related to this, the coupling of this economic concentration of wealth with the political system already concerns me. AI datacenters already represent a substantial fraction of US economic growth,44

44 需要明确的是,我并不认为实际的 AI 生产力已经对美国经济增长贡献了很大一部分。相反,我认为数据中心支出代表了由预期投资驱动的增长,这相当于市场预期未来 AI 驱动的经济增长并据此进行投资。

44 To be clear, I do not think actual AI productivity is yet responsible for a substantial fraction of US economic growth. Rather, I think the datacenter spending represents growth caused by anticipatory investment that amounts to the market expecting _future_ AI-driven economic growth and investing accordingly.

从而将大型科技公司(日益专注于 AI 或 AI 基础设施)的财务利益与政府的政治利益紧密捆绑在一起,这可能产生不正当的激励。我们已经看到这一点:科技公司不愿批评美国政府,而政府支持极端的 AI 放松管制政策。

and are thus strongly tying together the financial interests of large tech companies (which are increasingly focused on either AI or AI infrastructure) and the political interests of the government in a way that can produce perverse incentives. We already see this through the reluctance of tech companies to criticize the US government, and the government’s support for extreme anti-regulatory policies on AI.

防御措施 Defenses

对此我们能做些什么?首先,也是最明显的,公司应该简单地选择不参与其中。Anthropic 一直努力成为政策参与者而非政治参与者,并在任何政府下保持我们真实的观点。我们公开支持明智的 AI 监管和符合公共利益的出口管制,即使这些与政府政策相悖。

What can be done about this? First, and most obviously, companies should simply choose not to be part of it. Anthropic has always strived to be a policy actor and not a political one, and to maintain our authentic views whatever the administration. We’ve spoken up in favor of sensible AI regulation and export controls that are in the public interest, even when these are at odds with government policy.45

当我们同意政府的观点时,我们会说出来,并寻找双方都支持且真正对世界有益的政策共识点。我们的目标是成为诚实的中间人,而不是任何特定政党的支持者或反对者。

45 When we agree with the administration, we say so, and we look for points of agreement where mutually supported policies are genuinely good for the world. We are aiming to be honest brokers rather than backers or opponents of any given political party.

许多人告诉我,我们应该停止这样做,因为这可能导致不利的待遇,但在我们这样做的这一年里,Anthropic 的估值增长了超过 6 倍,这在我们这样的商业规模上几乎是前所未有的增长。

Many people have told me that we should stop doing this, that it could lead to unfavorable treatment, but in the year we’ve been doing it, Anthropic’s valuation has increased by over 6x, an almost unprecedented jump at our commercial scale.

其次,AI 行业需要与政府建立更健康的关系——一种基于实质性政策参与而非政治结盟的关系。我们选择参与政策实质而非政治,有时被解读为战术失误或未能“察言观色”,而非原则性决定,这种框架让我担忧。在一个健康的民主国家,公司应该能够为了政策本身而倡导良好政策。与此相关,公众对 AI 的反弹正在酝酿:这可能是一种纠正,但目前缺乏焦点。许多反弹针对的并非真正的问题(如数据中心的用水),并提出了无法解决真正担忧的解决方案(如数据中心禁令或设计不当的财富税)。值得关注的潜在问题是确保 AI 的发展对公共利益负责,不被任何特定的政治或商业联盟所俘获,将公众讨论集中于此似乎很重要。

Second, the AI industry needs a healthier relationship with government—one based on substantive policy engagement rather than political alignment. Our choice to engage on policy substance rather than politics is sometimes read as a tactical error or failure to “read the room” rather than a principled decision, and that framing concerns me. In a healthy democracy, companies should be able to advocate for good policy for its own sake. Related to this, a public backlash against AI is brewing: this could be a corrective, but it’s currently unfocused. Much of it targets issues that aren’t actually problems (like datacenterwater usage) and proposes solutions (like datacenter bans or poorly designed wealth taxes) that wouldn’t address the real concerns. The underlying issue that deserves attention is ensuring that AI development remains accountable to the public interest, not captured by any particular political or commercial alliance, and it seems important to focus the public discussion there.

第三,我在本节前面描述的宏观经济干预措施,以及私人慈善事业的复兴,可以帮助平衡经济天平,同时解决就业替代和经济权力集中问题。我们应该借鉴我们国家的历史:即使在镀金时代,像洛克菲勒和卡内基这样的实业家也对社会负有强烈的责任感,认为社会对他们的成功贡献巨大,他们需要回馈。这种精神如今似乎越来越缺失,我认为这是摆脱这一经济困境的重要途径。那些处于 AI 经济繁荣前沿的人应该愿意放弃他们的财富和权力。

Third, the macroeconomic interventions I described earlier in this section, as well as a resurgence of private philanthropy, can help to balance the economic scales, addressing both the job displacement and concentration of economic power problems at once. We should look to the history of our country here: even in the Gilded Age, industrialists such as Rockefeller and Carnegie felt a strong obligation to society at large, a feeling that society had contributed enormously to their success and they needed to give back. That spirit seems to be increasingly missing today, and I think it is a large part of the way out of this economic dilemma. Those who are at the forefront of AI’s economic boom should be willing to give away both their wealth and their power.

间接影响 Indirect effects

最后一节是未知未知的杂项,特别是那些可能因 AI 的积极进展以及由此带来的科学和技术整体加速而产生的间接问题。假设我们解决了迄今为止描述的所有风险,并开始收获 AI 的益处。我们很可能会迎来“一个世纪的科学和经济进步压缩到十年内”,这对世界来说将是巨大的积极影响,但随后我们将不得不应对这种快速进步带来的问题,而且这些问题可能会迅速出现。我们还可能遇到其他因 AI 进展间接产生的风险,这些风险很难提前预料。

This last section is a catchall for unknown unknowns, particularly things that could go wrong as an indirect result of positive advances in AI and the resulting acceleration of science and technology in general. Suppose we address all the risks described so far, and begin to reap the benefits of AI. We will likely get a “century of scientific and economic progress compressed into a decade,” and this will be hugely positive for the world, but we will then have to contend with the problems that arise from this rapid rate of progress, and those problems may come at us fast. We may also encounter other risks that occur indirectly as a consequence of AI progress and are hard to anticipate in advance.

由于未知未知的性质,不可能列出详尽的清单,但我将列出三个可能的担忧作为我们应关注的示例:

By the nature of unknown unknowns it is impossible to make an exhaustive list, but I’ll list three possible concerns as illustrative examples for what we should be watching for:

* **生物学的快速进步。** 如果我们在几年内取得一个世纪的医学进步,我们可能会大幅延长人类寿命,并且有可能获得激进的能力,比如提高人类智力或彻底改变人类生物学。这些将是可能性的巨大变化,而且发生得非常快。如果负责任地进行,它们可能是积极的(正如我在《Machines of Loving Grace》中所希望的),但总存在它们出大问题的风险——例如,如果让人类更聪明的努力也使他们更不稳定或更追求权力。还有“上传”或“全脑仿真”的问题,即在软件中实例化的数字人类思维,这有朝一日可能帮助人类超越其物理限制,但也带来了我认为令人不安的风险。

* Rapid advances in biology.If we do get a century of medical progress in a few years, it is possible that we will greatly increase the human lifespan, and there is a chance we also gain radical capabilities like the ability to increase human intelligence or radically modify human biology. Those would be big changes in what is possible, happening very quickly. They could be positive if responsibly done (which is my hope, as described in _Machines of Loving Grace_), but there is always a risk they go very wrong—for example, if efforts to make humans smarter also make them more unstable or power-seeking. There is also the issue of “uploads” or “whole brain emulation,” digital human minds instantiated in software, which might someday help humanity transcend its physical limitations, but which also carry risks I find disquieting.

* **AI 以不健康的方式改变人类生活。** 一个拥有数十亿在所有方面都比人类聪明得多的智能体的世界将是一个非常奇怪的世界。即使 AI 不主动攻击人类(第 1 节),也没有被国家明确用于压迫或控制(第 3 节),在正常商业激励和名义上自愿的交易中,仍有很多可能出错的地方。我们在关于 AI 精神病、AI 导致人们自杀以及关于与 AI 建立浪漫关系的担忧中看到了早期迹象。例如,强大的 AI 能否发明某种新宗教并让数百万人皈依?大多数人是否会以某种方式“上瘾”于与 AI 的互动?人们是否会被 AI 系统“操纵”,即 AI 基本上监视他们的一举一动,并随时告诉他们该做什么和说什么,导致一种“美好”但缺乏自由或任何成就感的生活?如果我和《黑镜》的创作者坐下来集思广益,不难想出几十个这样的场景。我认为这指出了改进 Claude 的 Constitution 之类事情的重要性,这超出了防止第 1 节问题所需的范畴。确保 AI 模型真正将用户的长期利益放在心上,以深思熟虑的人会认可的方式,而不是某种微妙扭曲的方式,似乎至关重要。

* AI changes human life in an unhealthy way.A world with billions of intelligences that are much smarter than humans at everything is going to be a very weird world to live in. Even if AI doesn’t actively aim to attack humans (Section 1), and isn’t explicitly used for oppression or control by states (Section 3), there is a lot that could go wrong short of this, via normal business incentives and nominally consensual transactions. We see early hints of this in the concerns about AI psychosis, AI driving people to suicide, and concerns about romantic relationships with AIs. As an example, could powerful AIs invent some new religion and convert millions of people to it? Could most people end up “addicted” in some way to AI interactions? Could people end up being “puppeted” by AI systems, where an AI essentially watches their every move and tells them exactly what to do and say at all times, leading to a “good” life but one that lacks freedom or any pride of accomplishment? It would not be hard to generate dozens of these scenarios if I sat down with the creator of _Black Mirror_ and tried to brainstorm them. I think this points to the importance of things like improving Claude’s Constitution, over and above what is necessary for preventing the issues in Section 1. Making sure that AI models _really_ have their users’ long-term interests at heart, in a way thoughtful people would endorse rather than in some subtly distorted way, seems critical.

* **人类目的。** 这与前一点相关,但与其说是关于人类与 AI 系统的具体互动,不如说是关于在拥有强大 AI 的世界中人类生活总体上如何变化。人类能否在这样的世界中找到目的和意义?我认为这是一个态度问题:正如我在《Machines of Loving Grace》中所说,我认为人类目的并不取决于在某方面成为世界最佳,人类可以通过他们热爱的故事和项目在很长一段时间内找到目的。我们只需要打破经济价值产生与自我价值和意义之间的联系。但这是社会必须做出的转变,而且总存在我们处理不当的风险。

* Human purpose.This is related to the previous point, but it’s not so much about specific human interactions with AI systems as it is about how human life changes in general in a world with powerful AI. Will humans be able to find purpose and meaning in such a world? I think this is a matter of attitude: as I said in _Machines of Loving Grace_, I think human purpose does not depend on being the best in the world at something, and humans can find purpose even over very long periods of time through stories and projects that they love. We simply need to break the link between the generation of economic value and self-worth and meaning. But that is a transition society has to make, and there is always the risk we don’t handle it well.

对于所有这些潜在问题,我的希望是,在一个拥有强大 AI 且我们信任它不会杀死我们、它也不是压迫性政府的工具、并且它真正为我们工作的世界里,我们可以利用 AI 本身来预测和预防这些问题。但这并不能保证——就像所有其他风险一样,这是我们必须谨慎处理的事情。

My hope with all of these potential problems is that in a world with powerful AI that we trust not to kill us, that is not the tool of an oppressive government, and that is genuinely working on our behalf, we can use AI itself to anticipate and prevent these problems. But that is not guaranteed—like all of the other risks, it is something we have to handle with care.

人类的考验 Humanity’s test

阅读这篇文章可能会让人觉得我们正处在一个令人生畏的境地。我写这篇文章时确实感到艰巨,这与《有爱的机器》形成对比,后者感觉像是为多年来在我脑海中回响的无比美妙的音乐赋予形式和结构。而关于这个局面,确实有很多困难之处。人工智能从多个方向给人类带来威胁,不同危险之间存在着真正的张力,如果我们不能极其谨慎地走钢丝,缓解其中一些风险可能会使其他风险变得更糟。

Reading this essay may give the impression that we are in a daunting situation. I certainly found it daunting to write, in contrast with _Machines of Loving Grace,_ which felt like giving form and structure to surpassingly beautiful music that had been echoing in my head for years. And there is much about the situation that genuinely is hard. AI brings threats to humanity from multiple directions, and there is genuine tension between the different dangers, where mitigating some of them risks making others worse if we do not thread the needle extremely carefully.

花时间谨慎地构建人工智能系统,使其不会自主威胁人类,这与民主国家需要保持对威权国家的领先地位、不被其征服之间存在真正的张力。但反过来,对抗专制所必需的人工智能工具,如果使用过度,也可能被用于在国内制造暴政。人工智能驱动的恐怖主义可能通过滥用生物学导致数百万人死亡,但对这一风险的过度反应可能使我们走向专制监控国家的道路。人工智能的劳动力和经济集中效应,除了本身是严重问题外,还可能迫使我们在一个公众愤怒甚至可能发生内乱的环境中面对其他问题,而不是能够诉诸我们本性中更善良的天使。最重要的是,风险的数量之多,包括未知风险,以及需要同时应对所有风险,构成了人类必须穿越的令人生畏的考验。

Taking time to carefully build AI systems so they do not autonomously threaten humanity is in genuine tension with the need for democratic nations to stay ahead of authoritarian nations and not be subjugated by them. But in turn, the same AI-enabled tools that are necessary to fight autocracies can, if taken too far, be turned inward to create tyranny in our own countries. AI-driven terrorism could kill millions through the misuse of biology, but an overreaction to this risk could lead us down the road to an autocratic surveillance state. The labor and economic concentration effects of AI, in addition to being grave problems in their own right, may force us to face the other problems in an environment of public anger and perhaps even civil unrest, rather than being able to call on the better angels of our nature. Above all, the sheer _number_ of risks, including unknown ones, and the need to deal with all of them at once, creates an intimidating gauntlet that humanity must run.

此外,过去几年应该已经清楚,停止甚至大幅减缓这项技术的想法从根本上说是站不住脚的。构建强大人工智能系统的公式极其简单,以至于几乎可以说它从数据和原始算力的正确组合中自发涌现。它的创造可能早在人类发明晶体管时就已不可避免,或者甚至更早,当我们第一次学会控制火时。如果一家公司不构建它,其他公司也会以几乎相同的速度构建。如果民主国家的所有公司通过相互协议或监管法令停止或减缓发展,那么威权国家只会继续前进。鉴于这项技术巨大的经济和军事价值,加上缺乏任何有意义的执行机制,我看不出我们如何能说服他们停止。

Furthermore, the last few years should make clear that the idea of stopping or even substantially slowing the technology is fundamentally untenable. The formula for building powerful AI systems is incredibly simple, so much so that it can almost be said to emerge spontaneously from the right combination of data and raw computation. Its creation was probably inevitable the instant humanity invented the transistor, or arguably even earlier when we first learned to control fire. If one company does not build it, others will do so nearly as fast. If all companies in democratic countries stopped or slowed development, by mutual agreement or regulatory decree, then authoritarian countries would simply keep going. Given the incredible economic and military value of the technology, together with the lack of any meaningful enforcement mechanism, I don’t see how we could possibly convince them to stop.

我确实看到了一条与现实主义地缘政治观相容的、略微减缓人工智能发展的路径。这条路径涉及通过剥夺威权国家构建强大人工智能所需的资源,即芯片和半导体制造设备,来将其向强大人工智能的进程延缓几年。

I do see a path to a _slight_ moderation in AI development that is compatible with a realist view of geopolitics). That path involves slowing down the march of autocracies towards powerful AI for a few years by denying them the resources they need to build it,46

我不认为超过几年是可能的:在更长的时间尺度上,他们将构建自己的芯片。

46 I don’t think anything more than a few years is possible: on longer timescales, they will build their own chips.

这反过来为民主国家提供了一个缓冲期,他们可以“花费”这个缓冲期来更谨慎地构建强大的人工智能,更多地关注其风险,同时仍然以足够快的速度前进,以轻松击败威权国家。民主国家内部人工智能公司之间的竞争随后可以在共同法律框架的庇护下,通过行业标准和监管的混合来处理。

namely chips and semiconductor manufacturing equipment. This in turn gives democratic countries a buffer that they can “spend” on building powerful AI more carefully, with more attention to its risks, while still proceeding fast enough to comfortably beat the autocracies. The race between AI companies within democracies can then be handled under the umbrella of a common legal framework, via a mixture of industry standards and regulation.

Anthropic 非常努力地倡导这条路径,推动芯片出口管制和对人工智能的明智监管,但即使是这些看似常识性的提议,也基本上被美国的政策制定者拒绝了(美国是最需要这些措施的国家)。人工智能可以带来如此巨大的利润——每年数万亿美元——以至于即使是最简单的措施也难以克服人工智能固有的政治经济学。这就是陷阱:人工智能如此强大,如此璀璨的奖品,以至于人类文明很难对其施加任何限制。

Anthropic has advocated very hard for this path, by pushing for chip export controls and judicious regulation of AI, but even these seemingly common-sense proposals have largely been rejected by policymakers in the United States (which is the country where it’s most important to have them). There is so much money to be made with AI—literally trillions of dollars per year—that even the simplest measures are finding it difficult to overcome the political economy inherent in AI. This is the trap: AI is so powerful, such a glittering prize, that it is very difficult for human civilization to impose any restraints on it at all.

我可以想象,正如萨根在《接触》中所写的那样,同样的故事在成千上万个世界上演。一个物种获得知觉,学会使用工具,开始技术的指数级上升,面临工业化和核武器的危机,如果它幸存下来,就会面临最艰难也是最终的挑战:当它学会如何将沙子塑造成会思考的机器时。我们能否通过这一考验,继续建设《有爱的机器》中描述的美丽社会,还是屈服于奴役和毁灭,将取决于我们作为一个物种的品格和决心,我们的精神和灵魂。

I can imagine, as Sagan did in _Contact_, that this same story plays out on thousands of worlds. A species gains sentience, learns to use tools, begins the exponential ascent of technology, faces the crises of industrialization and nuclear weapons, and if it survives those, confronts the hardest and final challenge when it learns how to shape sand into machines that think. Whether we survive that test and go on to build the beautiful society described in _Machines of Loving Grace_, or succumb to slavery and destruction, will depend on our character and our determination as a species, our spirit and our soul.

尽管存在许多障碍,我相信人类内心有力量通过这一考验。我受到鼓舞和激励,成千上万的研究人员将职业生涯奉献给了帮助我们理解和引导人工智能模型,以及塑造这些模型的品格和构成。我认为现在有很大的机会,这些努力能够及时结出果实,产生重要影响。我受到鼓舞,至少有一些公司表示,他们将承担有意义的商业成本,以阻止他们的模型助长生物恐怖主义的威胁。我受到鼓舞,少数勇敢的人抵制了当前的政治风向,通过了立法,为人工智能系统设置了第一批早期合理的护栏。我受到鼓舞,公众理解人工智能带有风险,并希望这些风险得到解决。我受到鼓舞,世界各地自由的不屈精神以及抵抗任何地方暴政的决心。

Despite the many obstacles, I believe humanity has the strength inside itself to pass this test. I am encouraged and inspired by the thousands of researchers who have devoted their careers to helping us understand and steer AI models, and to shaping the character and constitution of these models. I think there is now a good chance that those efforts bear fruit in time to matter. I am encouraged that at least some companies have stated they’ll pay meaningful commercial costs to block their models from contributing to the threat of bioterrorism. I am encouraged that a few brave people have resisted the prevailing political winds and passedlegislation that puts the first early seeds of sensible guardrails on AI systems. I am encouraged that the public understands that AI carries risks and wants those risks addressed. I am encouraged by the indomitable spirit of freedom around the world and the determination to resist tyranny wherever it occurs.

但如果我们想要成功,我们需要加大努力。第一步是让最接近技术的人简单地告诉人类所处的真实情况,我一直试图这样做;在这篇文章中,我更加明确和紧迫地这样做。下一步将是说服世界的思想家、政策制定者、公司和公民,让他们认识到这个问题的紧迫性和压倒一切的重要性——值得为此投入思考和政治资本,而不是每天主导新闻的数千个其他问题。然后,将是一个需要勇气的时刻,足够多的人能够逆流而上,坚持原则,即使面临对其经济利益和个人安全的威胁。

But we will need to step up our efforts if we want to succeed. The first step is for those closest to the technology to simply tell the truth about the situation humanity is in, which I have always tried to do; I’m doing so more explicitly and with greater urgency with this essay. The next step will be convincing the world’s thinkers, policymakers, companies, and citizens of the imminence and overriding importance of this issue—that it is worth expending thought and political capital on this in comparison to the thousands of other issues that dominate the news every day. Then there will be a time for courage, for enough people to buck the prevailing trends and stand on principle, even in the face of threats to their economic interests and personal safety.

未来的岁月将异常艰难,对我们的要求超出我们所能给予的。但在我作为研究者、领导者和公民的岁月里,我看到了足够的勇气和高尚,让我相信我们能够获胜——当被置于最黑暗的环境中时,人类有一种方式,似乎在最后一刻,聚集起取胜所需的力量和智慧。我们没有时间可以浪费。

The years in front of us will be impossibly hard, asking more of us than we think we can give. But in my time as a researcher, leader, and citizen, I have seen enough courage and nobility to believe that we can win—that when put in the darkest circumstances, humanity has a way of gathering, seemingly at the last minute, the strength and wisdom needed to prevail. We have no time to lose.

我要感谢 Erik Brynjolfsson、Ben Buchanan、Mariano-Florentino Cuéllar、Allan Dafoe、Kevin Esvelt、Nick Beckstead、Richard Fontaine、Jim McClave 以及 Anthropic 的许多员工,感谢他们对本文草稿提出的有益意见。

I would like to thank Erik Brynjolfsson, Ben Buchanan, Mariano-Florentino Cuéllar, Allan Dafoe, Kevin Esvelt, Nick Beckstead, Richard Fontaine, Jim McClave, and very many of the staff at Anthropic for their helpful comments on drafts of this essay.

互动版:图/公式 + 针对本篇提问 →