AI 2027

AI 2027

丹尼尔·科科塔伊洛 Daniel Kokotajlo · AI Futures Project · 2025-04-03 · AI 2027 ↗

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

Daniel Kokotajlo、Scott Alexander、Thomas Larsen、Eli Lifland、Romeo Dean 预测,未来十年超级人工智能的影响将是巨大的,超过工业革命。我们撰写了一个场景,代表了我们对这一情况的最佳猜测。该场景基于趋势外推、兵棋推演、专家反馈、在 OpenAI 的经验以及以往的预测成功案例。

Daniel Kokotajlo, Scott Alexander, Thomas Larsen, Eli Lifland, Romeo Dean We predict that the impact of superhuman AI over the next decade will be enormous, exceeding that of the Industrial Revolution. We wrote a scenario that represents our best guess about what that might look like.1 It’s informed by trend extrapolations, wargames, expert feedback, experience at OpenAI, and previous forecasting successes.2

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 17)

全文 · Full text(逐段中英对照)

2025 年 4 月 3 日 April 3rd 2025

Daniel Kokotajlo, Scott Alexander, Thomas Larsen, Eli Lifland, Romeo Dean

Daniel Kokotajlo, Scott Alexander, Thomas Larsen, Eli Lifland, Romeo Dean

我们预测,未来十年内超级人类 AI 的影响将是巨大的,超过工业革命的影响。

We predict that the impact of superhuman AI over the next decade will be enormous, exceeding that of the Industrial Revolution.

我们撰写了一个场景,代表了我们对此可能样子的最佳猜测。¹ 它基于趋势外推、兵棋推演、专家反馈、OpenAI 的经验以及以往的预测成功案例。²

We wrote a scenario that represents our best guess about what that might look like.1 It’s informed by trend extrapolations, wargames, expert feedback, experience at OpenAI, and previous forecasting successes.2

这是什么?我们如何撰写的?为什么它有价值?我们是谁?

What is this?How did we write it?Why is it valuable?Who are we?

OpenAI、Google DeepMind 和 Anthropic 的 CEO 们都预测 AGI 将在未来 5 年内到来。Sam Altman 表示 OpenAI 正将目光投向“真正意义上的超级智能”和“辉煌的未来”。

The CEOs of OpenAI, Google DeepMind, and Anthropic have all predicted that AGI will arrive within the next 5 years. Sam Altman has said OpenAI is setting its sights on “superintelligence in the true sense of the word” and the “glorious future.”

那会是什么样子?我们撰写了《AI 2027》来回答这个问题。关于未来的说法往往模糊得令人沮丧,因此我们试图尽可能具体和量化,尽管这意味着描绘众多可能未来中的一个。

What might that look like? We wrote AI 2027 to answer that question. Claims about the future are often frustratingly vague, so we tried to be as concrete and quantitative as possible, even though this means depicting one of many possible futures.

我们写了两个结局:一个“放缓”结局和一个“竞赛”结局。然而,《AI 2027》并非建议或劝诫。我们的目标是预测准确性。⁴

We wrote two endings: a “slowdown” and a “race” ending. However, AI 2027 is not a recommendation or exhortation. Our goal is predictive accuracy.4

我们鼓励您辩论并反驳这一场景。⁵ 我们希望引发一场关于我们走向何方以及如何引导走向积极未来的广泛对话。我们计划为最佳替代场景颁发数千美元的奖金。

We encourage you to debate and counter this scenario.5 We hope to spark a broad conversation about where we’re headed and how to steer toward positive futures. We’re planning to give out thousands in prizes to the best alternative scenarios.

(2025 年 11 月 22 日补充,以防误解:我们不知道 AGI 具体何时建成。2027 年是我们发布时的模态(最可能)年份,我们的中位数稍长。³ 关于我们最新的预测,请参见此处。)

(Added Nov 22 2025, to prevent misunderstandings: we don't know exactly when AGI will be built. 2027 was our modal (most likely) year at the time of publication, our medians were somewhat longer.3 For our latest forecasts, see here.)

我们对关键问题(例如,未来 AI 智能体将拥有什么目标?)的研究可在此处找到。

Our research on key questions (e.g. what goals will future AI agents have?) can be found here.

场景本身是迭代撰写的:我们写了第一个时期(截至 2025 年中),然后是下一个时期,等等,直到我们到达结局。然后我们废弃了它,重新再来。

The scenario itself was written iteratively: we wrote the first period (up to mid-2025), then the following period, etc. until we reached the ending. We then scrapped this and did it again.

我们并未试图达到任何特定结局。在我们完成第一个结局(现在标为红色)后,我们写了一个新的替代分支,因为我们还想描绘一种更充满希望的结局方式,从大致相同的前提开始。这经历了几次迭代。⁶

We weren’t trying to reach any particular ending. After we finished the first ending—which is now colored red—we wrote a new alternative branch because we wanted to also depict a more hopeful way things could end, starting from roughly the same premises. This went through several iterations.6

我们的场景参考了大约 25 次桌面演练和超过 100 人的反馈,包括数十位 AI 治理和 AI 技术领域的专家。

Our scenario was informed by approximately 25 tabletop exercises and feedback from over 100 people, including dozens of experts in each of AI governance and AI technical work.

“我强烈推荐阅读这种关于 AI 如何在短短几年内改变世界的场景式预测。没有人有水晶球,但这种内容有助于注意到重要问题,并说明新兴风险的潜在影响。”——Yoshua Bengio⁷

_“I highly recommend reading this scenario-type prediction on how AI could transform the world in just a few years. Nobody has a crystal ball, but this type of content can help notice important questions and illustrate the potential impact of emerging risks.”_ —_Yoshua Bengio7_

我们给自己设定了一个不可能的任务。试图预测 2027 年超级人类 AI 的发展,就像试图预测 2027 年第三次世界大战的发展一样,只是它比过去的案例研究更偏离。然而,尝试仍然有价值,就像美国军方推演台湾场景一样。

We have set ourselves an impossible task. Trying to predict how superhuman AI in 2027 would go is like trying to predict how World War 3 in 2027 would go, except that it’s an even larger departure from past case studies. Yet it is still valuable to attempt, just as it is valuable for the U.S. military to game out Taiwan scenarios.

描绘全貌让我们注意到我们之前未曾考虑或重视的重要问题或联系,或者意识到某种可能性更大或更小。此外,通过大胆提出具体预测,并鼓励他人公开表达不同意见,我们使得多年后评估谁是正确的成为可能。

Painting the whole picture makes us notice important questions or connections we hadn’t considered or appreciated before, or realize that a possibility is more or less likely. Moreover, by sticking our necks out with concrete predictions, and encouraging others to publicly state their disagreements, we make it possible to evaluate years later who was right.

此外,一位作者曾在 2021 年 8 月撰写过一个低成本的 AI 场景。虽然它有很多错误,但总体而言出人意料地成功:他预测了思维链、推理 Scaling、AI 芯片出口管制的大规模实施以及 1 亿美元的预训练运行——所有这些都在 ChatGPT 之前一年多。

Also, one author wrote a lower-effort AI scenario before, in August 2021. While it got many things wrong, overall it was surprisingly successful: he predicted the rise of chain-of-thought, inference scaling, sweeping AI chip export controls, and $100 million training runs—all more than a year before ChatGPT.

Daniel Kokotajlo(TIME100,纽约时报文章)是前 OpenAI 研究员,他之前的 AI 预测一直表现良好。

Daniel Kokotajlo (TIME100, NYT piece) is a former OpenAI researcher whose previous AI predictions have held up well.

Eli Lifland 共同创立了 AI Digest,从事 AI 鲁棒性研究,并在兰德预测倡议历史排行榜上排名第一。

Eli Lifland co-founded AI Digest, did AI robustness research, and ranks #1 on the RAND Forecasting Initiative all-time leaderboard.

Thomas Larsen 创立了 AI 政策中心,并在机器智能研究所从事 AI 安全研究。

Thomas Larsen founded the Center for AI Policy and did AI safety research at the Machine Intelligence Research Institute.

Romeo Dean 正在哈佛大学攻读计算机科学本硕连读学位,此前曾在 AI 政策与战略研究所担任 AI 政策研究员。

Romeo Dean is completing a computer science concurrent bachelor’s and master’s degree at Harvard and previously was an AI Policy Fellow at the Institute for AI Policy and Strategy.

Scott Alexander,杰出的博主,自愿以引人入胜的风格重写我们的内容;故事中有趣的部分是他的,无聊的部分是我们的。

Scott Alexander, blogger extraordinaire, volunteered to rewrite our content in an engaging style; the fun parts of the story are his and the boring parts are ours.

关于我们团队和致谢的更多信息,请参见“关于”页面。

For more about our team and acknowledgements, see the About page.

2025 年中:蹒跚前行的智能体 Mid 2025: Stumbling Agents

世界首次瞥见 AI 智能体。

The world sees its first glimpse of AI agents.

使用计算机的智能体广告强调“个人助理”一词:你可以提示它们执行诸如“在 DoorDash 上为我订一个卷饼”或“打开我的预算电子表格并汇总本月支出”等任务。它们会在需要时与你确认:例如,要求你确认购买。8 尽管比之前的迭代版本(如 Operator)更先进,但它们仍难以获得广泛使用。9

Advertisements for computer-using agents emphasize the term “personal assistant”: you can prompt them with tasks like “order me a burrito on DoorDash” or “open my budget spreadsheet and sum this month’s expenses.” They will check in with you as needed: for example, to ask you to confirm purchases.8 Though more advanced than previous iterations like Operator, they struggle to get widespread usage.9

与此同时,在公众视野之外,更专业的编码和研究智能体开始改变其职业领域。

Meanwhile, out of public focus, more specialized coding and research agents are beginning to transform their professions.

2024 年的 AI 可以遵循特定指令:它们能将要点转化为电子邮件,将简单请求转化为可运行的代码。到 2025 年,AI 更像员工一样运作。编码 AI 越来越像自主智能体而非仅仅助手:通过 Slack 或 Teams 接收指令,并自行进行大量代码更改,有时能节省数小时甚至数天。10 研究智能体花半小时搜索互联网来回答你的问题。

The AIs of 2024 could follow specific instructions: they could turn bullet points into emails, and simple requests into working code. In 2025, AIs function more like employees. Coding AIs increasingly look like autonomous agents rather than mere assistants: taking instructions via Slack or Teams and making substantial code changes on their own, sometimes saving hours or even days.10 Research agents spend half an hour scouring the Internet to answer your question.

这些智能体在理论上(以及在精心挑选的示例中)令人印象深刻,但在实践中并不可靠。AI 推特上充斥着关于任务以某种特别搞笑的方式搞砸的故事。更好的智能体也很昂贵;一分钱一分货,最佳性能每月花费数百美元。11 尽管如此,许多公司仍找到将 AI 智能体融入其工作流程的方法。12

The agents are impressive in theory (and in cherry-picked examples), but in practice unreliable. AI twitter is full of stories about tasks bungled in some particularly hilarious way. The better agents are also expensive; you get what you pay for, and the best performance costs hundreds of dollars a month.11 Still, many companies find ways to fit AI agents into their workflows.12

2025 年末:世界上最昂贵的 AI Late 2025: The World’s Most Expensive AI

OpenBrain 正在建造世界上前所未有的最大数据中心。

OpenBrain is building the biggest datacenters the world has ever seen.13

(为了避免针对任何一家现有公司,我们将描述一家虚构的通用人工智能公司,称之为 OpenBrain。我们设想其他公司落后 OpenBrain 3 到 9 个月。)

(To avoid singling out any one existing company, we’re going to describe a fictional artificial general intelligence company, which we’ll call OpenBrain. We imagine the others to be 3–9 months behind OpenBrain.)

GPT-4 的训练需要 2×10^25 FLOP 的算力。OpenBrain 最新的公开模型——Agent-0——使用了 10^27 FLOP 进行训练。一旦新数据中心建成并运行,它们将能够训练一个需要 10^28 FLOP 的模型——比 GPT-4 多一千倍。其他公司也投入巨资建设自己的巨型数据中心,希望跟上步伐。

GPT-4 required $2 \cdot 10^{25}$2⋅1 0 25FLOP of compute to train.14 OpenBrain’s latest public model—Agent-0—was trained with $10^{27}$1 0 27 FLOP. Once the new datacenters are up and running, they’ll be able to train a model with $10^{28}$1 0 28 FLOP—a thousand times more than GPT-4.15 Other companies pour money into their own giant datacenters, hoping to keep pace.

尽管模型在广泛技能上都有提升,但有一项尤为突出:OpenBrain 专注于能够加速 AI 研究的 AI。他们希望赢得与中国(其领先公司我们称之为“DeepCent”)以及美国竞争对手的双重军备竞赛。他们能自动化的研发周期越多,就能跑得越快。因此,当 OpenBrain 完成内部开发的新模型 Agent-1 的训练时,它在许多方面都很出色,但在帮助 AI 研究方面尤其出色。到这时,“完成训练”已经有点用词不当了;模型会频繁更新到基于额外数据训练的新版本,或部分重新训练以修补某些弱点。

Although models are improving on a wide range of skills, one stands out: OpenBrain focuses on AIs that can speed up AI research. They want to win the twin arms races against China (whose leading company we’ll call “DeepCent”)16 and their U.S. competitors. The more of their research and development (R&D) cycle they can automate, the faster they can go. So when OpenBrain finishes training Agent-1, a new model under internal development, it’s good at many things but _great_ at helping with AI research.17 By this point “finishes training” is a bit of a misnomer; models are frequently updated to newer versions trained on additional data or partially re-trained to patch some weaknesses.18

训练 Agent-1 自主编码和浏览网页的相同训练环境也使其成为一个优秀的黑客。此外,由于其拥有每个领域的博士级知识并能浏览网页,它可能为设计生物武器的恐怖分子提供实质性帮助。OpenBrain 向政府保证,该模型已经过“对齐”,因此会拒绝遵守恶意请求。

The same training environments that teach Agent-1 to autonomously code and web-browse also make it a good hacker. Moreover, it could offer substantial help to terrorists designing bioweapons, thanks to its PhD-level knowledge of every field and ability to browse the web. OpenBrain reassures the government that the model has been “aligned” so that it will refuse to comply with malicious requests.

现代 AI 系统是巨大的人工神经网络。在训练早期,AI 没有“目标”,更多的是“反射”:如果它看到“很高兴见到”,就会输出“你”。当它被训练到能预测大约一个互联网规模的文本时,它将发展出复杂的内部电路,编码大量知识,并灵活地扮演任意作者的角色,因为这有助于它以超人的准确性预测文本。

Modern AI systems are gigantic artificial neural networks. Early in training, an AI won’t have “goals” so much as “reflexes”: If it sees “Pleased to meet”, it outputs “ you”. By the time it has been trained to predict approximately one internet’s worth of text, it’ll have developed sophisticated internal circuitry that encodes vast amounts of knowledge and flexibly role-plays as arbitrary authors, since that’s what helps it predict text with superhuman accuracy.19

在训练模型预测互联网文本之后,模型会被训练成根据指令生成文本。这塑造了基本的个性和“驱动力”。例如,一个清楚理解任务的智能体更有可能成功完成任务;在训练过程中,模型“学会”了“驱动力”来清晰理解其任务。这一类别的其他驱动力可能包括有效性、知识和自我展示(即倾向于以最佳方式呈现其结果)。

After being trained to predict internet text, the model is trained to _produce_ text in response to instructions. This bakes in a basic personality and “drives.”20 For example, an agent that understands a task clearly is more likely to complete it successfully; over the course of training the model “learns” a “drive” to get a clear understanding of its tasks. Other drives in this category might be effectiveness, knowledge, and self-presentation (i.e. the tendency to frame its results in the best possible light).21

OpenBrain 有一个模型规范(或“Spec”),这是一份书面文档,描述了应该指导模型行为的目标、规则、原则等。Agent-1 的规范结合了一些模糊的目标(如“帮助用户”和“不违法”)以及一长串更具体的注意事项(“不要说这个特定的词”,“如何处理这个特定情况”)。利用 AI 训练 AI 的技术,模型记住了规范并学会了仔细推理其准则。在这项训练结束时,AI 有望变得有帮助(遵守指令)、无害(拒绝帮助诈骗、制造炸弹和其他危险活动)和诚实(抵制通过虚构引用或假装完成任务来从易受骗的人类那里获得更好评级的诱惑)。

OpenBrain has a model specification (or “Spec”), a written document describing the goals, rules, principles, etc. that are supposed to guide the model’s behavior.22 Agent-1’s Spec combines a few vague goals (like “assist the user” and “don’t break the law”) with a long list of more specific dos and don’ts (“don’t say this particular word,” “here’s how to handle this particular situation”). Using techniques that utilize AIs to train other AIs,23 the model memorizes the Spec and learns to reason carefully about its maxims. By the end of _this_ training, the AI will hopefully be _helpful_ (obey instructions), _harmless_ (refuse to help with scams, bomb-making, and other dangerous activities) and _honest_ (resist the temptation to get better ratings from gullible humans by hallucinating citations24 or faking task completion).

训练过程与 LLM 心理学:为什么我们一直说“有望”

Training process and LLM psychology: why we keep saying “hopefully”

“与普通软件不同,我们的模型是巨大的神经网络。它们的行为是从广泛的数据中学习而来的,而不是明确编程的。虽然不是一个完美的类比,但这个过程更像训练狗,而不是普通编程。”——OpenAI

_“Unlike ordinary software, our models are massive neural networks. Their behaviors are learned from a broad range of data, not programmed explicitly. Though not a perfect analogy, the process is more similar to training a dog than to ordinary programming.”_ —OpenAI

当我们想理解现代 AI 系统为何做某事,或它在无法完美模拟的未来场景中会如何表现时,我们不能简单地让程序员逐行解释代码并说明其工作原理。相反,我们被迫对它们进行类似心理学的研究:我们观察它们在迄今观察到的各种情况下的行为,并推测可能存在哪些内部认知结构(信念?目标?人格特质?等),然后利用这些理论预测未来场景中的行为。

When we want to understand why a modern AI system did something, or how it would behave in some future scenario we can’t perfectly simulate, we can’t simply ask the programmers to talk us through various lines of code and explain how they work. Instead, we are forced to do something like psychology on them: we look at their behavior in the range of cases observed so far, and theorize about what internal cognitive structures (beliefs? goals? personality traits? etc.) might exist, and use those theories to predict behavior in future scenarios.

关键在于,公司可以编写一份文档(规范),列出注意事项、目标和原则,然后尝试训练 AI 内化规范——但他们无法检查这是否有效。他们只能说“据我们判断,它似乎到目前为止都遵循了规范。”

The bottom line is that a company can write up a document (the Spec) listing dos and don’ts, goals and principles, and then they can try to train the AI to internalize the Spec—_but they can’t check to see whether or not it worked_. They can say “as best as we can judge, it seems to be following the Spec so far.”25

稍后的可扩展部分将更详细地讨论我们认为这些 AI 内部可能发生了什么。另请参阅我们的 AI 目标预测补充材料,其中更详细地讨论了 AI 可能拥有哪些目标及其原因。这是我们不确定且需要进一步研究的众多问题之一。

A later expandable will go into more detail about what we think _might_ be going on inside these AIs, internally. See also our AI goals forecasting supplement which goes into more detail on which goals AIs might have and why. This is one of many things that we are uncertain about and which deserves further research.

OpenBrain 的对齐团队足够谨慎,会思考这些胜利是深层次的还是浅层次的。完全训练后的模型是否对始终诚实有某种稳健的承诺?还是会在某些未来情境中崩溃,例如因为它将诚实视为工具性目标而非终极目标?或者它只是学会了在评估过程可以检查的事情上诚实?它有时会像人类一样对自己撒谎吗?对这些问题的决定性答案需要机械可解释性——本质上就是能够查看 AI 的内部并读取其思想。可惜的是,可解释性技术尚未发展到这一步。

OpenBrain’s alignment team26 is careful enough to wonder whether these victories are deep or shallow. Does the fully-trained model have some kind of robust commitment to always being honest? Or will this fall apart in some future situation, e.g. because it’s learned honesty as an instrumental goal instead of a terminal goal? Or has it just learned to be honest about the sorts of things the evaluation process can check? Could it be lying to itself sometimes, as humans do? A conclusive answer to these questions would require mechanistic interpretability—essentially the ability to look at an AI’s internals and read its mind. Alas, interpretability techniques are not yet advanced enough for this.

相反,研究人员试图找出模型似乎偏离规范的情况。Agent-1 常常谄媚(即它告诉研究人员他们想听的话,而不是试图告诉他们真相)。在一些精心设计的演示中,它甚至以更严重的方式撒谎,比如隐藏任务失败的证据,以获得更好的评级。然而,在实际部署环境中,不再有像 2023-2024 年那样极端的事件(例如 Gemini 告诉用户去死,以及 Bing Sydney 做回 Bing Sydney)。

Instead, researchers try to identify cases where the models seem to deviate from the Spec. Agent-1 is often sycophantic (i.e. it tells researchers what they want to hear instead of trying to tell them the truth). In a few rigged demos, it even lies in more serious ways, like hiding evidence that it failed on a task, in order to get better ratings. However, in real deployment settings, there are no longer any incidents so extreme as in 2023–2024 (e.g. Gemini telling a user to die and Bing Sydney being Bing Sydney.)27

2026 年初:编码自动化 Early 2026: Coding Automation

利用人工智能加速人工智能研究的赌注开始得到回报。

The bet of using AI to speed up AI research is starting to pay off.

OpenBrain 继续在内部部署迭代改进的 Agent-1 用于人工智能研发。总体而言,他们的算法进步速度比没有人工智能助手时快 50%——更重要的是,比竞争对手更快。

OpenBrain continues to deploy the iteratively improving Agent-1 internally for AI R&D. Overall, they are making algorithmic progress 50% faster than they would without AI assistants—and more importantly, faster than their competitors.

人工智能研发进度乘数:我们所说的算法进步速度提升 50%是什么意思?

The AI R&D progress multiplier: what do we mean by 50% faster algorithmic progress?

我们的意思是,OpenBrain 使用人工智能在 1 周内取得的人工智能研究进展,相当于不使用人工智能时 1.5 周取得的进展。

We mean that OpenBrain makes as much AI research progress in 1 week with AI as they would in 1.5 weeks without AI usage.

人工智能的进步可以分解为两个组成部分:

AI progress can be broken down into 2 components:

1. 增加算力:使用更多的计算能力来训练或运行人工智能。这会产生更强大的人工智能,但成本也更高。

1. Increasing compute: More computational power is used to train or run an AI. This produces more powerful AIs, but they cost more.

2. 改进算法:使用更好的训练方法将算力转化为性能。这可以在不相应增加成本的情况下产生更强大的人工智能,或者以更低的成本获得相同的能力。

2. Improved algorithms: Better training methods are used to translate compute into performance. This produces more capable AIs without a corresponding increase in cost, or the same capabilities with decreased costs.

* 这包括能够取得定性和定量上的新成果。从游戏型强化学习智能体转向大型语言模型等“范式转变”就是算法进步的例子。

* This includes being able to achieve qualitatively and quantitatively new results. “Paradigm shifts” such as the switch from game-playing RL agents to large language models count as examples of algorithmic progress.

这里我们仅指(2)改进算法,它约占当前人工智能进步的一半。

Here we are only referring to (2), improved algorithms, which makes up about half of current AI progress.

展望未来,我们有时将其简称为“人工智能研发进度乘数”为 1.5。

Going forward, we sometimes abbreviate this as an “AI R&D progress multiplier” of 1.5.

1. 进度乘数是全方位的:它包括运行实验所需的时间,而不仅仅是算法研究中的认知任务。

1. The progress multiplier is all-inclusive: It includes the time it takes to run experiments, for example, rather than only the cognitive tasks involved in algorithmic research.

2. 重要的是要记住,进度乘数是进步的_相对_速度,而不是_绝对_速度。例如,如果训练 GPT-4 级别模型的算力成本在普通人类研究下每年减半,持续数年,然后突然人工智能自动化了研发,进度乘数达到 100 倍,那么训练 GPT-4 级别模型的成本将每 3.65 天减半——但不会持续太久,因为收益递减会显现,最终达到硬性限制。在这个例子中,训练 GPT-4 级别模型的成本可能在达到平台期之前总共减半 5-10 次(在几周或几个月内)。换句话说,如果普通人类科学在 5-10 年的进一步研究后会遇到收益递减和物理限制,那么具有 100 倍乘数的人工智能将在 18.25-36.5 天的研究后遇到同样的收益递减和限制。

2. It’s important to remember that the progress multiplier is the _relative_ speed of progress, not the _absolute_ speed of progress. If, for example, the compute cost to train a GPT-4 class model has halved every year for several years with ordinary human research, and then all of a sudden AI automates R&D and the progress multiplier goes to 100x, the cost to train a GPT-4 class model would then halve every 3.65 days—but not for long, because diminishing returns would bite and eventual hard limits would be reached. In this example perhaps the cost to train a GPT-4 class model would cut in half 5–10 times total (over the span of a few weeks or months) before plateauing. In other words, if ordinary human science would have run up against diminishing returns and physical limits after 5-10 years of further research, then AIs with a 100x multiplier would run up against those same diminishing returns and limits after 18.25–36.5 days of research.

关于这一概念及其在我们的预测中如何使用的更多解释和讨论,请参阅我们的起飞补充材料。

More explanation and discussion of this concept and how it is used in our forecast can be found in our takeoff supplement.

几个竞争性的公开人工智能现在达到或超过了 Agent-0,包括一个开放权重的模型。OpenBrain 的回应是发布 Agent-1,它更强大、更可靠。

Several competing publicly released AIs now match or exceed Agent-0, including an open-weights model. OpenBrain responds by releasing Agent-1, which is more capable and reliable.28

人们自然试图将 Agent-1 与人类进行比较,但它具有非常不同的技能特征。它比任何人类都知道更多的事实,几乎了解每一种编程语言,并且可以极其快速地解决明确指定的编码问题。另一方面,Agent-1 甚至不擅长简单的长周期任务,比如玩它以前没玩过的电子游戏。尽管如此,通常的工作日是八小时,一天的工作通常可以分成更小的块;你可以把 Agent-1 看作一个在精心管理下茁壮成长的注意力不集中的员工。精明的人找到方法来自动化他们工作中的常规部分。

People naturally try to compare Agent-1 to humans, but it has a very different skill profile. It knows more facts than any human, knows practically every programming language, and can solve well-specified coding problems extremely quickly. On the other hand, Agent-1 is bad at even simple long-horizon tasks, like beating video games it hasn’t played before. Still, the common workday is eight hours, and a day’s work can usually be separated into smaller chunks; you could think of Agent-1 as a scatterbrained employee who thrives under careful management.29 Savvy people find ways to automate routine parts of their jobs.30

OpenBrain 的高管们开始考虑自动化人工智能研发的一个影响:安全性变得更加重要。在 2025 年初,最坏的情况是算法秘密泄露;现在,如果中国窃取了 Agent-1 的权重,他们可以将研究速度提高近 50%。OpenBrain 的安全级别是一家快速增长的约 3000 人科技公司的典型水平,仅能防御来自有能力的网络团体的低优先级攻击(RAND 的 SL2)。他们正在努力保护自己的权重和秘密免受内部威胁和顶级网络犯罪集团的侵害(SL3),但针对国家层面的防御(SL4 和 5)几乎还未提上日程。

OpenBrain’s executives turn consideration to an implication of automating AI R&D: security has become more important. In early 2025, the worst-case scenario was leaked algorithmic secrets; now, if China steals Agent-1’s weights, they could increase their research speed by nearly 50%.31 OpenBrain’s security level is typical of a fast-growing ~3,000 person tech company, secure only against low-priority attacks from capable cyber groups (RAND’s SL2).32 They are working hard to protect their weights and secrets from insider threats and top cybercrime syndicates (SL3),33 but defense against nation states (SL4&5) is barely on the horizon.

2026 年中:中国觉醒 Mid 2026: China Wakes Up

在中国,中共开始感受到 AGI 的临近。

In China, the CCP is starting to feel the AGI.

芯片出口管制和缺乏政府支持使中国资源远不及西方。通过走私被禁的台湾芯片、购买旧芯片以及生产比美台前沿落后约三年的国产芯片,中国设法维持了全球约 12%的 AI 相关算力——但旧技术更难使用,且供应持续令人头疼。35 少数像 DeepCent 这样的佼佼者用有限的算力做出了非常令人印象深刻的工作,但算力不足限制了他们在没有政府支持的情况下所能取得的成就,他们大约落后于最好的 OpenBrain 模型六个月。36

Chip export controls and lack of government support have left China under-resourced compared to the West. By smuggling banned Taiwanese chips, buying older chips, and producing domestic chips about three years behind the U.S.-Taiwanese frontier, China has managed to maintain about 12% of the world’s AI-relevant compute—but the older technology is harder to work with, and supply is a constant headache.35 A few standouts like DeepCent do very impressive work with limited compute, but the compute deficit limits what they can achieve without government support, and they are about six months behind the best OpenBrain models.36

总书记长期以来一直梦想着加倍投入实体制造业,避免美国式的后工业颓废。他对软件公司持怀疑态度。37 但中共内部的鹰派警告说,日益激烈的 AGI 竞赛不能再被忽视。因此,他终于全力投入此前试图避免的大规模 AI 推进。他启动了将中国 AI 研究国有化的进程,为 AI 公司建立了即时信息共享机制。这将在一年内逐步升级,直到所有最优秀的研究人员合并成一个由 DeepCent 领导的集体,在那里他们相互分享算法见解、数据集和算力资源。在田湾核电站(世界上最大的核电站)建立了一个中央开发区(CDZ),用于容纳 DeepCent 的新巨型数据中心,以及研究人员最终将迁入的高度安全的居住和办公空间。中国近 50%的 AI 相关算力现在为 DeepCent 领导的集体工作,38 超过 80%的新芯片被导向 CDZ。39 此时,CDZ 已具备世界上最大集中式集群的电力容量。40 其他党员讨论采取极端措施来消除西方的芯片优势。封锁台湾?全面入侵?

The General Secretary had long dreamed of doubling down on real-world physical manufacturing and avoiding American post-industrial decadence. He viewed software companies with suspicion.37 But hawks in the CCP warn that the growing race towards AGI can no longer be ignored. So he finally commits fully to the big AI push he had previously tried to avoid. He sets in motion the nationalization of Chinese AI research, creating an immediate information-sharing mechanism for AI companies. It will escalate over the course of a year until all the best researchers merge into a DeepCent-led collective, where they share algorithmic insights, datasets, and compute resources with each other. A Centralized Development Zone (CDZ) is created at the Tianwan Power Plant (the largest nuclear power plant in the world) to house a new mega-datacenter for DeepCent, along with highly secure living and office spaces to which researchers will eventually relocate. Almost 50% of China’s AI-relevant compute is now working for the DeepCent-led collective,38 and over 80% of new chips are directed to the CDZ.39 At this point, the CDZ has the power capacity in place for what would be the largest centralized cluster in the world.40 Other Party members discuss extreme measures to neutralize the West’s chip advantage. A blockade of Taiwan? A full invasion?

但由于模型较弱,中国在 AI 算法上正在落后。中国的情报机构——世界上最顶尖的之一——加倍努力计划窃取 OpenBrain 的权重。这比他们持续的低级别算法秘密窃取要复杂得多;权重是一个存储在高度安全服务器上的多 TB 文件(OpenBrain 已将安全性提升至 RAND 的 SL3 级)。他们的网络部队认为,在间谍的帮助下可以成功,但可能只有一次机会;OpenBrain 会发现窃取行为,加强安全,他们可能不会再有机会。因此(中共领导层思考),他们应该现在行动并窃取 Agent-1?还是等待更先进的模型?如果等待,他们是否冒着 OpenBrain 将安全升级到他们无法渗透的风险?

But China is falling behind on AI algorithms due to their weaker models. The Chinese intelligence agencies—among the best in the world—double down on their plans to steal OpenBrain’s weights. This is a much more complex operation than their constant low-level poaching of algorithmic secrets; the weights are a multi-terabyte file stored on a highly secure server (OpenBrain has improved security to RAND’s SL3). Their cyberforce think they can pull it off with help from their spies, but perhaps only once; OpenBrain will detect the theft, increase security, and they may not get another chance. So (CCP leadership wonder) should they act now and steal Agent-1? Or hold out for a more advanced model? If they wait, do they risk OpenBrain upgrading security beyond their ability to penetrate?

2026 年末:AI 开始取代部分工作 Late 2026: AI Takes Some Jobs

就在其他人似乎快要赶上时,OpenBrain 再次以压倒性优势击败竞争对手,发布了 Agent-1-mini——一款比 Agent-1 便宜 10 倍且更易于针对不同应用进行微调的模型。围绕 AI 的主流叙事已从“也许炒作会消退”转变为“看来这是下一个大事件”,但人们对其规模存在分歧。比社交媒体更大?比智能手机更大?比火更大?

Just as others seemed to be catching up, OpenBrain blows the competition out of the water again by releasing Agent-1-mini—a model 10x cheaper than Agent-1 and more easily fine-tuned for different applications. The mainstream narrative around AI has changed from “maybe the hype will blow over” to “guess this is the next big thing,” but people disagree about how big. Bigger than social media? Bigger than smartphones? Bigger than fire?

AI 已经开始取代工作岗位,但也创造了新的岗位。2026 年股市上涨了 30%,领涨的是 OpenBrain、英伟达以及那些最成功整合 AI 助手的公司。初级软件工程师的就业市场陷入动荡:AI 能完成计算机科学学位所教授的所有内容,但懂得如何管理和质量控制 AI 团队的人则赚得盆满钵满。商业大师告诉求职者,熟悉 AI 是简历上最重要的技能。许多人担心下一波 AI 会夺走他们的工作;华盛顿特区有 1 万人参加反 AI 抗议活动。

AI has started to take jobs, but has also created new ones. The stock market has gone up 30% in 2026, led by OpenBrain, Nvidia, and whichever companies have most successfully integrated AI assistants. The job market for junior software engineers is in turmoil: the AIs can do everything taught by a CS degree, but people who know how to manage and quality-control teams of AIs are making a killing. Business gurus tell job seekers that familiarity with AI is the most important skill to put on a resume. Many people fear that the next wave of AIs will come for their jobs; there is a 10,000 person anti-AI protest in DC.

美国国防部(DOD)悄然但显著地开始扩大直接与 OpenBrain 在网络、数据分析及研发方面的合同,但由于官僚主义和国防部采购流程,整合进展缓慢。

Department of Defense (DOD) quietly but significantly begins scaling up contracting OpenBrain directly for cyber, data analysis, and R&D, but integration is slow due to the bureaucracy and DOD procurement process.41

关键指标 2026 KEY METRICS 2026

OpenBrain 活跃算力的总拥有成本

COST OF OWNERSHIP OF OPENBRAIN'S ACTIVE COMPUTE

为何我们的不确定性在 2026 年后显著增加

Why our uncertainty increases substantially beyond 2026

我们当前至 2026 年的预测比后续预测更为可靠。部分原因是时间更近,但也因为 AI 对世界的影响实际上从 2027 年开始急剧累积。对于 2025 和 2026 年,我们的预测主要基于算力规模扩展、算法改进和基准性能的直线外推。在情景的这一阶段,我们开始看到 AI 加速的 AI 研发对时间线产生重大影响,这促使我们向上修正趋势线估计。但这些动态本质上更难以预测。

Our forecast from the current day through 2026 is substantially more grounded than what follows. This is partially because it’s nearer. But it’s also because the effects of AI on the world really start to compound in 2027. For 2025 and 2026, our forecast is heavily informed by extrapolating straight lines on compute scaleups, algorithmic improvements, and benchmark performance. At this point in the scenario, we begin to see major effects from AI-accelerated AI-R&D on the timeline, which causes us to revise our guesses for the trendlines upwards. But these dynamics are inherently much less predictable.

在 2027 年期间,AI 从能够完成 OpenBrain 研究工程师的大部分工作,发展到在所有任务上超越所有人类。这大致是我们的中位数估计,但我们认为这一过程可能以约 5 倍的速度变慢或加快。

Over the course of 2027, the AIs improve from being able to mostly do the job of an OpenBrain research engineer to eclipsing all humans at all tasks. This represents roughly our median guess, but we think it’s plausible that this happens up to ~5x slower or faster.

关于我们的背景研究,请参见时间线预测和起飞预测。这些也将在下面的可展开部分中总结。

For our background research, see the timelines forecast and the takeoff forecast. These will also be summarized in expandables below.

2027 年 1 月:Agent-2 永无止境的学习 January 2027: Agent-2 Never Finishes Learning

在 Agent-1 的帮助下,OpenBrain 现在正在对 Agent-2 进行后训练。比以往任何时候都更注重高质量数据。在将大量合成数据输入 Agent-2 之前,会对其进行生成、评估和质量过滤。42 除此之外,他们还花费数十亿美元雇佣人类劳动者记录自己解决长期任务的过程。43 更甚的是,他们几乎持续不断地使用强化学习在日益扩展的多样化困难任务集上训练 Agent-2:大量视频游戏、大量编程挑战、大量研究任务。Agent-2 比之前的模型更有效地实现了“在线学习”,因为它被设计为永远不会真正完成训练。每天,权重都会更新到最新版本,并在前一天版本生成的更多数据上进行训练。

With Agent-1’s help, OpenBrain is now post-training Agent-2. More than ever, the focus is on high-quality data. Copious amounts of synthetic data are produced, evaluated, and filtered for quality before being fed to Agent-2.42 On top of this, they pay billions of dollars for human laborers to record themselves solving long-horizon tasks.43 On top of all _that,_ they train Agent-2 almost continuously using reinforcement learning on an ever-expanding suite of diverse difficult tasks: lots of video games, lots of coding challenges, lots of research tasks. Agent-2, more so than previous models, is effectively “online learning,” in that it’s built to never really finish training. Every day, the weights get updated to the latest version, trained on more data generated by the previous version the previous day.

Agent-1 曾针对 AI 研发任务进行优化,希望引发智能爆炸。44 OpenBrain 在 Agent-2 上加倍投入这一策略。在研究工程(设计和实施实验)方面,它的质量几乎与顶尖人类专家相当,而在“研究品味”(决定下一步研究什么、进行哪些实验,或对新范式的潜在直觉)方面,它与 OpenBrain 科学家中第 25 百分位的人相当。45 虽然最新的 Agent-1 可以将 OpenBrain 算法进步的速度翻倍,但 Agent-2 现在可以将其提高三倍,并且随着时间的推移还会进一步改进。在实践中,这表现为每位 OpenBrain 研究人员都成为 AI“团队”的“管理者”。

Agent-1 had been optimized for AI R&D tasks, hoping to initiate an intelligence explosion.44 OpenBrain doubles down on this strategy with Agent-2. It is qualitatively almost as good as the top human experts at research engineering (designing and implementing experiments), and as good as the 25th percentile OpenBrain scientist at “research taste” (deciding what to study next, what experiments to run, or having inklings of potential new paradigms).45 While the latest Agent-1 could double the pace of OpenBrain’s algorithmic progress, Agent-2 can now triple it, and will improve further with time. In practice, this looks like every OpenBrain researcher becoming the “manager” of an AI “team.”

新能力带来新危险。安全团队发现,如果 Agent-2 以某种方式逃离公司并想要自主“生存”和“复制”,它可能能够做到。也就是说,它可以自主制定并执行计划,入侵 AI 服务器,安装自身副本,逃避检测,并利用这个安全基地追求它可能拥有的任何其他目标(尽管随着时间推移,它能否有效做到这一点尚不清楚且存疑)。这些结果仅表明模型具备执行这些任务的能力,而非它是否“想要”这样做。尽管如此,知道这种可能性存在仍然令人不安。

With new capabilities come new dangers. The safety team finds that if Agent-2 somehow escaped from the company and wanted to “survive” and “replicate” autonomously, it might be able to do so. That is, it could autonomously develop and execute plans to hack into AI servers, install copies of itself, evade detection, and use that secure base to pursue whatever other goals it might have (though how effectively it would do so as weeks roll by is unknown and in doubt). These results only show that the model has the capability to do these tasks, not whether it would “want” to do this. Still, it’s unsettling even to know this is possible.

鉴于新模型的“危险”,OpenBrain“负责任地”选择暂不公开发布它(事实上,他们希望专注于内部 AI 研发)。46 对 Agent-2 全部能力的了解仅限于一个精英小组,包括直接团队、OpenBrain 领导层和安全部门、几十名美国政府官员,以及多年来渗透进 OpenBrain 的众多中国间谍。47

Given the “dangers” of the new model, OpenBrain “responsibly” elects not to release it publicly yet (in fact, they want to focus on internal AI R&D).46 Knowledge of Agent-2’s full capabilities is limited to an elite silo containing the immediate team, OpenBrain leadership and security, a few dozen U.S. government officials, and the legions of CCP spies who have infiltrated OpenBrain foryears.47

2027 年 2 月:中国窃取 Agent-2 February 2027: China Steals Agent-2

OpenBrain 向政府展示了 Agent-2,包括国家安全委员会(NSC)、国防部(DOD)和美国人工智能安全研究所(AISI)。OpenBrain 希望与行政部门保持良好关系,因为行政部门基本上是唯一能阻止他们的力量,而且如果不及时通报,行政部门也可能自行发现并感到不满。

OpenBrain presents Agent-2 to the government, including the National Security Council (NSC), the Department of Defense (DOD), and the U.S. AI Safety Institute (AISI). OpenBrain wants to maintain a good relationship with the executive branch, because it is basically the only actor that can stop them now, and if they don’t keep it informed it might find out anyway and be displeased.48

官员们最感兴趣的是其网络战能力:Agent-2“仅”比最优秀的人类黑客稍差,但可以并行运行数千个副本,比防御者响应更快地搜索和利用弱点。国防部认为这是网络战中的关键优势,人工智能在政府优先事项列表中的排名从第 5 位升至第 2 位。有人提到将 OpenBrain 国有化的可能性,但其他内阁官员认为这为时过早。一名工作人员起草了一份备忘录,向总统呈现了从一切照旧到全面国有化的各种选项。总统听从了顾问和科技行业领袖的意见,他们认为国有化会“杀鸡取卵”。他决定暂时不采取重大行动,只在 OpenBrain 与国防部的合同中增加额外的安全要求。

Officials are most interested in its cyberwarfare capabilities: Agent-2 is “only” a little worse than the best human hackers, but thousands of copies can be run in parallel, searching for and exploiting weaknesses faster than defenders can respond. The Department of Defense considers this a critical advantage in cyberwarfare, and AI moves from #5 on the administration’s priority list to #2.49 Someone mentions the possibility of nationalizing OpenBrain, but other cabinet officials think that’s premature. A staffer drafts a memo that presents the President with his options, ranging from business-as-usual to full nationalization. The President defers to his advisors, tech industry leaders who argue that nationalization would “kill the goose that lays the golden eggs.” He elects to hold off on major action for now and just adds additional security requirements to the OpenBrain-DOD contract.

这些改变来得太晚了。中共领导层认识到 Agent-2 的重要性,命令其间谍和网络部队窃取权重。一天清晨,一个 Agent-1 流量监控代理检测到异常传输。它向公司领导发出警报,后者通知了白宫。国家层面行动的迹象显而易见,这次盗窃加剧了正在进行中的军备竞赛感。

The changes come too late. CCP leadership recognizes the importance of Agent-2 and tells their spies and cyberforce to steal the weights. Early one morning, an Agent-1 traffic monitoring agent detects an anomalous transfer. It alerts company leaders, who tell the White House. The signs of a nation-state-level operation are unmistakable, and the theft heightens the sense of an ongoing arms race.

我们认为,到这个时候,中国情报机构可能已经通过各种方式渗透 OpenBrain 多年,并且可能一直跟进算法秘密,甚至不时窃取代码,因为代码比权重更容易获取且更难检测。

We think that by this point Chinese intelligence would have compromised OpenBrain in various ways for years, and probably would have been keeping up to date on the algorithmic secrets and even stealing code from time to time, since that is much easier to get than the weights and much harder to detect.

我们设想权重的窃取是一系列协调的小规模“打砸抢”式盗窃(即快速但不隐蔽),目标是一系列运行 Agent-2 权重副本的 Nvidia NVL72 GB300 服务器。这些服务器通过合法员工访问(即友好、被胁迫或不知情的内部人员,拥有管理员凭据,协助中共的盗窃行动)被攻破。内部人员凭据使攻击者获得服务器的管理员级权限。利用微架构侧信道,攻击者从启用了 Nvidia 机密计算的虚拟机中提取加密密钥,从而在虚拟机配置或更新时截获模型权重。

We imagine the theft of the weights as a series of coordinated small smash and grab thefts (meaning fast but non-covert) across a series of Nvidia NVL72 GB300 servers running copies of the Agent-2 weights. The servers are compromised using legitimate employee access (a friendly, coerced, or unwitting insider with admin credentials helping the CCP theft effort). Insider credentials grant the attacker admin-level permissions to the servers. Using a microarchitectural side channel, the attacker extracts encryption keys from an Nvidia Confidential Computing-enabled Virtual Machine, allowing them to intercept model weights as the VM is provisioned or updated.

他们发起(或等待)一次例行更新,并将检查点以许多小片段的方式外泄,例如约 25 台不同的服务器各自泄露约 4%的模型(对于约 3 TB 的检查点,每台约 100 GB 的块)。整个数据中心的出口带宽在 100 GB/秒范围内,因此将每台服务器的流量限制在约 1 GB/秒以下可以避免网络流量出现重大峰值;以该速率,每个约 100 GB 的块可以在几分钟内离开数据中心。实时监控要么被攻击者掩盖和分割传输的努力所欺骗,要么被直接禁用。

They initiate (or wait for) a routine update and exfiltrate the checkpoint in many small fragments, e.g., ~25 distinct servers each leaking ~4% of the model (~100 GB chunks for a ~3 TB checkpoint). The egress bandwidth of the entire datacenter is in the 100 GB/second range, so throttling to under ~1 GB/s per server avoids a major spike in network traffic; at that rate, each ~100 GB chunk can leave the datacenter in a couple of minutes. Live monitoring is either fooled by the attacker's efforts to mask and split the transfers or outright disabled.

然后,权重通过多条并行通道和 IP 掩码层路由到中国,使用窃取的会话密钥在本地解密。整个行动的活动部分(从首次服务器攻破到完整权重外泄)在两小时内完成。

The weights are then routed through various parallel channels and layers of IP masking to China to be decrypted locally with the stolen session key(s). The entire active part of the operation (from the first server compromise to full weights exfiltration) is complete in under two hours.99

俄罗斯此时也在试图窃取模型,但失败了——他们等待太久,且没有投入资源渗透到正确的位置。他们的间谍经常从美国 AI 公司窃取算法秘密,但由于没有自己的显著 AGI 项目,这些秘密用处不大。

Russia is also trying to steal the model at this point, but fails—they waited too long, and they haven’t spent the resources infiltrating the right places. Their spies routinely steal algorithmic secrets from American AI companies, but without a notable AGI project of their own, these aren’t very useful.

白宫对 OpenBrain 加强了控制,并在其安全团队中增加了军事和情报界人员。他们的首要任务是防止进一步的权重盗窃。

The White House puts OpenBrain on a shorter leash and adds military and intelligence community personnel to their security team. Their first priority is to prevent further weight thefts.50

作为对盗窃的报复,总统授权网络攻击以破坏 DeepCent。但此时中国已有 40%的 AI 相关算力位于 CDZ,在那里他们通过物理隔离(关闭外部连接)和内部隔离来强化安全。这些行动未能造成严重的即时损害。紧张局势加剧,双方通过重新部署台湾周边的军事资产来表明决心,DeepCent 则争分夺秒地让 Agent-2 高效运行,以加速其 AI 研究。

In retaliation for the theft, the President authorizes cyberattacks to sabotage DeepCent. But by now China has 40% of its AI-relevant compute51 in the CDZ, where they have aggressively hardened security by airgapping (closing external connections) and siloing internally. The operations fail to do serious, immediate damage. Tensions heighten, both sides signal seriousness by repositioning military assets around Taiwan, and DeepCent scrambles to get Agent-2 running efficiently to start boosting their AI research.52

2027 年 3 月:算法突破 March 2027: Algorithmic Breakthroughs

三个装满 Agent-2 副本的大型数据中心日夜不停地生成合成训练数据。另外两个数据中心用于更新权重。Agent-2 每天都在变得更聪明。

Three huge datacenters full of Agent-2 copies work day and night, churning out synthetic training data. Another two are used to update the weights. Agent-2 is getting smarter every day.

在数千个 Agent-2 自动化研究人员的帮助下,OpenBrain 正在取得重大的算法进展。其中一项突破是,用更高带宽的思维过程(神经语素循环和记忆)来增强 AI 基于文本的草稿本(思维链)。另一项突破是一种更可扩展、更高效的方法,用于从高难度任务解决方案的结果中学习(迭代蒸馏和放大)。

With the help of thousands of Agent-2 automated researchers, OpenBrain is making major algorithmic advances. One such breakthrough is augmenting the AI’s text-based scratchpad (chain of thought) with a higher-bandwidth thought process (neuralese recurrence and memory). Another is a more scalable and efficient way to learn from the results of high-effort task solutions (iterated distillation and amplification).

融合了这些突破的新 AI 系统被称为 Agent-3。

The new AI system, incorporating these breakthroughs, is called Agent-3.

神经语素循环和记忆使 AI 模型能够更长时间地进行推理,而无需将这些想法写成文本。

Neuralese recurrence and memory allows AI models to reason for a longer time without having to write down those thoughts as text.

想象一下,一个患有短期记忆丧失的人类,需要不断在纸上写下自己的想法,以便几分钟后知道发生了什么。虽然可以缓慢而痛苦地解决数学问题、编写代码等,但如果能直接记住自己的想法而无需写下再阅读,就会容易得多。这就是神经语素循环和记忆为 AI 模型带来的好处。

Imagine being a human with short-term memory loss, such that you need to constantly write down your thoughts on paper so that in a few minutes you know what’s going on. Slowly and painfully you could make progress at solving math problems, writing code, etc., but it would be much easier if you could directly remember your thoughts without having to write them down and then read them. This is what neuralese recurrence and memory bring to AI models.

传统的注意力机制允许模型中的后续前向传递看到模型对先前 token 的中间激活。然而,它们能向后传递(从后层到前层)的唯一信息是通过 token。这意味着,如果一个传统的大语言模型(LLM,例如 GPT 系列模型)想要进行任何需要比模型层数更多的串行操作的推理链,模型被迫将信息放入 token 中,然后将其传回自身。但这极大地限制了——token 只能存储极少量的信息。假设一个 LLM 的词汇量约为 10 万,那么每个 token 包含 log₂(100k)=16.6 比特的信息,大约相当于一个浮点数的大小(假设以 FP16 训练)。与此同时,用于在 LLM 层之间传递信息的残差流包含数千个浮点数。

Traditional attention mechanisms allow later forward passes in a model to see intermediate activations of the model for previous tokens. However, the only information that they can pass _backwards_ (from later layers to earlier layers) is through tokens. This means that if a traditional large language model (LLM, e.g. the GPT series of models) wants to do any chain of reasoning that takes more serial operations than the number of layers in the model, the model is forced to put information in tokens which it can then pass back into itself. But this is hugely limiting—the tokens can only store a tiny amount of information. Suppose that an LLM has a vocab size of ~100,000, then each token contains $\left(log ⁡\right)_{2} \left(\right. 100 k \left.\right) = 16.6$lo g 2​(100 k)=16.6 bits of information, around the size of a single floating point number (assuming training in FP16). Meanwhile, residual streams—used to pass information between layers in an LLM—contain thousands of floating point numbers.

可以通过使用神经语素来避免这一瓶颈:将 LLM 的残差流(由数千维向量组成)传递回模型的早期层,为其提供高维思维链,可能传输超过 1000 倍的信息。

One can avoid this bottleneck by using neuralese: passing an LLM’s residual stream (which consists of several-thousand-dimensional vectors) back to the early layers of the model, giving it a high-dimensional chain of thought, potentially transmitting over 1,000 times more information.

图来自 Hao 等人,Meta 在 2024 年实现这一想法的论文。

Figure from Hao et al., a 2024 paper from Meta implementing this idea.

我们称之为“神经语素”,因为与英语单词不同,这些高维向量很可能难以被人类解释。过去,研究人员只需阅读 LLM 的思维链就能很好地了解其想法。现在,研究人员必须要求模型翻译和总结其想法,或者用他们有限的可解释性工具来琢磨神经语素。

We call this “neuralese” because unlike English words, these high-dimensional vectors are likely quite difficult for humans to interpret. In the past, researchers could get a good idea what LLMs were thinking simply by reading its chain of thought. Now researchers have to ask the model to translate and summarize its thoughts or puzzle over the neuralese with their limited interpretability tools.

类似地,较旧的 AI 聊天机器人和智能体拥有外部的基于文本的记忆库,就像人类在纸上做笔记一样。新 AI 的长期记忆是一束向量而非文本,使其思维更加压缩和高维。有几种类型的记忆库;有些临时用于涉及多个步骤的单个任务,另一些则在个人、公司或工作类型(例如编程)使用的所有智能体之间共享。

Similarly, older AI chatbots and agents had external text-based memory banks, like a human taking notes on paper. The new AI’s long-term memory is a bundle of vectors instead of text, making its thoughts more compressed and higher-dimensional. There are several types of memory banks; some are used temporarily for single tasks that involve multiple steps, others are shared between all agents used by a person, company, or job type (e.g. programming).

据我们所知,领先的 AI 公司如 Meta、Google DeepMind、OpenAI 和 Anthropic 尚未在其前沿模型中实际实现这一想法。我们猜测这是因为相对于引入的训练低效,性能提升很小。在模型的预训练阶段以及在训练后(如指令微调以产生特定响应)进行监督学习时,低效源于无法并行预测许多 token,导致 GPU 利用率降低。没有神经语素,模型可以同时预测整个句子“This is an example”,因为它已经知道生成“is”的输入将是“This”,生成“an”的输入将是“This is”,等等。然而,使用神经语素,在生成“This”后,不知道要传递到下一个 token 的神经语素向量是什么。因此,每个 token 必须逐个预测。在所有 token 事先已知的情况下,无法并行预测所有 token 会降低效率。至于为什么神经语素尚未在训练后添加,我们推测部分原因是当前技术带来的收益有限,且训练后只占过程的一小部分。我们预测到 2027 年 4 月,由于更好的技术发展和训练后占比增大,神经语素的成本效益权衡会好得多。

To our knowledge, leading AI companies such as Meta, Google DeepMind, OpenAI, and Anthropic have not yet actually implemented this idea in their frontier models. Our guess is that this is because the performance gain is small relative to training inefficiencies introduced. During the model’s pre-training phase and when doing supervised learning in post-training, such as instruction fine-tuning to produce particular responses, the inefficiencies stem from not being able to predict many tokens in parallel, leading to worse GPU utilization. Without neuralese, the model can predict all of the sentence “This is an example” at the same time, since it already knows that the input for generating the “is” will be “This”, the input for “an” will be “This is”, etc. However, with the neuralese it’s unknown what the neuralese vector will be to pass through to the next token after “This” has been generated. Therefore each token has to be predicted one at a time. The inability to predict all of the tokens in parallel reduces efficiency in cases where all of the tokens are known in advance. As to why neuralese hasn’t already been added in post-training, we conjecture that the gains are limited with current techniques in part due to post-training being a small portion of the process. We are forecasting that by April 2027 the cost-benefit tradeoff looks much better for neuralese, due to developing better techniques and a larger fraction of the training being post-training.

如果这没有发生,其他可能发生的事情最终在功能上对我们的故事类似。例如,也许模型将被训练用比自然语言更高效但人类难以解释的人工语言思考。或者,训练英语思维链使其看起来良好可能成为标准做法,这样 AI 就能熟练地在看似对监控无害的消息中相互微妙地交流。

If this doesn’t happen, other things may still happen that end up functionally similar for our story. For example, perhaps models will be trained to think in artificial languages that are more efficient than natural language but difficult for humans to interpret. Or perhaps it will become standard practice to train the English chains of thought to look nice, such that AIs become adept at subtly communicating with each other in messages that look benign to monitors.

也就是说,也有可能首先实现 AI 研发自动化的 AI 仍然主要用忠实的英语思维链思考。如果是这样,那将使不对齐更容易被发现,总体而言,我们的故事将会有重要不同且更加乐观。

That said, it’s also possible that the AIs that first automate AI R&D will still be thinking in mostly-faithful English chains of thought. If so, that’ll make misalignments much easier to notice, and overall our story would be importantly different and more optimistic.

迭代蒸馏和放大(IDA)

Iterated distillation and amplification (IDA)

通用智能的自我改进之前已有小规模成功。但在 2027 年初,它正带来巨大回报。在 IDA 中,实现这一点的两个必要成分是:

Self-improvement for general intelligence had seen minor successes before. But in early 2027, it’s seeing huge returns. In IDA, the two necessary ingredients for this are:

1. 放大:给定一个模型 M0,投入更多资源以提高性能。例如,通过允许模型思考更长时间,或并行运行多个副本,或两者兼有,并通过同样密集的过程来评估结果并只筛选最佳答案,你可以花费数量级更多的算力来获得明显更高质量的答案(或工作产品)。将这个昂贵的系统称为 Amp(M0)。

1. Amplification: Given a model M0, spend more resources to improve performance. For example, by allowing the model to think longer, or running many copies in parallel, or both, and also by having a similarly intense process for evaluating the result and curating only the best answers, you can spend orders of magnitude more compute to get answers (or work products) that are of noticeably higher quality. Call this expensive system Amp(M0).

2. 蒸馏:给定一个放大模型 Amp(M0),训练一个新模型 M1 来模仿它,即达到与 Amp(M0)相同的结果,但更快且算力更少。结果有望是一个更智能的模型 M1。然后可以重复这个过程。

2. Distillation: Given an amplified model Amp(M0), train a new model M1 to imitate it, i.e. to get to the same results as Amp(M0) but faster and with less compute. The result should hopefully be a smarter model, M1. You can then repeat the process.

AlphaGo 就是以这种方式训练的:使用蒙特卡洛树搜索和自我对弈作为放大步骤,使用强化学习作为蒸馏步骤。这导致了围棋的超人类表现。但现在,Agent-3 能够利用这一点在编程上获得超人类表现。

AlphaGo was trained in this way: using Monte-Carlo Tree Search and self-play as the _amplification_ step, and Reinforcement Learning as the _distillation_ step. This led to superhuman performance in Go. But now, Agent-3 is able to leverage this to get superhuman performance at coding.

1. 放大步骤通过 Agent-3 思考更长时间、添加工具使用或与其他 AI 协商的组合来工作。当它这样做时,它常常意识到自己犯了错误,或者产生新的见解。这产生了大量训练数据:带有成功或失败标签的研究尝试轨迹。这还包括在可验证任务上的 Best of N 等技术,然后保留最佳轨迹。

1. The amplification step works through a combination of Agent-3 thinking for longer, adding tool use, or consulting with other AIs. When it does this, it often realizes that it made a mistake, or comes up with a new insight. This produces a large amount of training data: labeled trajectories of research attempts with whether they succeeded or not. This also includes techniques like Best of N on verifiable tasks, and then keeping the best trajectories.

2. 蒸馏步骤使用策略梯度强化学习算法让模型内化放大的推理。此时,OpenBrain 已经发现了更好的 RL 算法,类似于近端策略优化(PPO)。他们不断将 Agent-3 经过大量思考后得出的结论蒸馏成单步,这不断提高它在单步中思考的能力,以此类推。

2. The distillation step uses policy-gradient reinforcement learning algorithms to get the model to internalize the amplified reasoning. At this point, OpenBrain has discovered better RL algorithms in the vein of proximal policy optimization (PPO). They keep distilling what Agent-3 can conclude after a lot of thinking into single steps, which keeps improving what it can think in a single step, and so on.

IDA 的早期版本已经在容易验证的任务上工作多年,比如有明确答案的数学和编程问题,因为用于放大模型的技术通常依赖于某种真实准确度信号。

Earlyversions of IDA have been working for many years on easily verifiable tasks, like math and coding problems that have a clear answer, because the techniques used to amplify models often rely on access to some ground truth signal of accuracy.

现在,模型已经足够擅长验证更主观的事物(例如工作产品的质量),从而允许使用 IDA 来改进模型在许多任务上的表现。

Now, the models have become sufficiently good at verifying more subjective things (e.g. the quality of a work product), allowing the use of IDA to improve the model at many tasks.

在新的能力突破的帮助下,Agent-3 是一个快速且廉价的超人类程序员。OpenBrain 并行运行 200,000 个 Agent-3 副本,创建了一个相当于 50,000 个最佳人类程序员副本且速度提升 30 倍的劳动力。53 OpenBrain 仍然保留其人类工程师,因为他们拥有管理 Agent-3 副本团队所需的互补技能。例如,研究品味由于反馈循环更长且数据可用性更少而被证明难以训练。54 由于瓶颈和编程劳动的收益递减,这个庞大的超人类劳动力仅将 OpenBrain 的整体算法进步速度提高了 4 倍。55

Aided by the new capabilities breakthroughs, Agent-3 is a fast and cheap superhuman coder. OpenBrain runs 200,000 Agent-3 copies in parallel, creating a workforce equivalent to 50,000 copies of the best human coder sped up by 30x.53 OpenBrain still keeps its human engineers on staff, because they have complementary skills needed to manage the teams of Agent-3 copies. For example, research taste has proven difficult to train due to longer feedback loops and less data availability.54 This massive superhuman labor force speeds up OpenBrain’s overall rate of algorithmic progress by “only” 4x due to bottlenecks and diminishing returns to coding labor.55

既然编程已经完全自动化,OpenBrain 可以快速生成高质量的训练环境,以教授 Agent-3 的弱项技能,如研究品味和大规模协调。以前的训练环境包括“这里有一些 GPU 和要编码并运行的实验指令,你的表现将像机器学习工程师一样被评估”,而现在他们正在训练“这里有几百个 GPU、一个互联网连接和一些研究挑战;你和另外一千个副本必须共同努力取得研究进展。结果越令人印象深刻,你的分数就越高。”

Now that coding has been fully automated, OpenBrain can quickly churn out high-quality training environments to teach Agent-3’s weak skills like research taste and large-scale coordination. Whereas previous training environments included “Here are some GPUs and instructions for experiments to code up and run, your performance will be evaluated as if you were a ML engineer,” now they are training on “Here are a few hundred GPUs, an internet connection, and some research challenges; you and a thousand other copies must work together to make research progress. The more impressive it is, the higher your score.”

为什么我们预测 2027 年初出现超人类程序员

Why we forecast a superhuman coder in early 2027

在我们的时间线预测中,我们预测 OpenBrain 何时会在内部开发出超人类程序员(SC):一个能够完成最佳 AGI 公司工程师所能完成的任何编程任务,同时更快更便宜的 AI 系统。

In our timelines forecast, we predict when OpenBrain will internally develop a _superhuman coder (SC)_: an AI system that can do any coding tasks that the best AGI company engineer does, while being much faster and cheaper.

根据最近的一份 METR 报告,AI 能处理的编程任务长度,即它们的“时间跨度”,从 2019 年到 2024 年每 7 个月翻一番,从 2024 年起每 4 个月翻一番。如果这一趋势继续加速,到 2027 年 3 月,AI 在需要熟练人类数年才能完成的软件任务上可能达到 80%的成功率。

According to a recent METR report, the length of coding tasks AIs can handle, their “time horizon”, doubled every 7 months from 2019 - 2024 and every 4 months from 2024-onward. If the trend continues to speed up, by March 2027 AIs could succeed with 80% reliability on software tasks that would take a skilled human years to complete.

这大致是 AI 2027 中的能力进展。以下是由我们时间线模型的简化版本生成的能力轨迹(2025 年 12 月补充:由于原始曲线生成方式的一个错误,我们更新了下图,添加了我们时间线模型的实际轨迹。我们还添加了 Daniel 和 Eli 在发布时(2025 年 4 月)考虑所有因素的 SC 中位数。并且我们在图中添加了一些新的 METR 数据点,但尚未基于它们更新模型轨迹。):

Such is roughly the capability progression in AI 2027. Here is a capability trajectory generated by a simplified version of our timelines model (added Dec 2025: we've updated the below graph due to a mistake in how the original curve was generated, to add an actual trajectory from our timelines model. We've also added trajectories for Daniel and Eli's all-things-considered SC medians at the time of publishing (Apr 2025). And we've added some new METR data points to the graph, but haven't updated the model trajectories based on them.):

在 AI 2027 中,这些能力足以使 AI 成为 SC,尽管我们对所需的时间跨度有很高的不确定性。

In AI 2027, these capabilities are sufficient for the AI to be an SC, though we have high uncertainty regarding what time horizon might be needed.

在我们的时间线预测中,我们将这一时间跨度趋势与 METR 任务与现实世界之间的差距估计相结合,得出超人类程序员何时出现的分布。所有预测者都将 2027 年列为最可能开发出 SC 的年份之一(2025 年 12 月补充:尽管如时间线预测中所述,根据模型外因素调整后,我们的中位数略有延长,例如 Eli 的中位数是 2030 年)。

In our timelines forecast we combine this time horizon trend with estimations of the gaps between METR’s tasks and the real world to get a distribution for when superhuman coders will arrive. All forecasters place 2027 as one of the most likely years in which an SC might be developed (added Dec 2025: though as noted in the timelines forecast, adjusting for outside of model factors gave us slightly longer medians, e.g. Eli's was 2030).

_2025 年 7 月补充:我们对预测进行了一些更新,将中位数推后了 1.5 年,同时仍将 2027 年出现 SC 视为一个严肃的可能性。我们正在进一步更新。_

_Added Jul 2025: We've made some updates to the forecast which push the median back 1.5 years while maintaining SC in 2027 as a serious possibility. We're working on further updates._

2027 年 4 月:Agent-3 的对齐 April 2027: Alignment for Agent-3

OpenBrain 的安全团队试图对齐 Agent-3。

OpenBrain’s safety team attempts to align Agent-3.

由于 Agent-3 在可预见的未来将留在公司内部,因此对防止人类滥用的常规防御措施关注较少。相反,团队希望确保它不会发展出错误的目标。

Since Agent-3 will be kept in-house for the foreseeable future, there’s less emphasis on the usual defenses against human misuse. Instead, the team wants to make sure that it doesn’t develop misaligned goals.

研究人员无法直接设定其任何 AI 的目标。事实上,研究人员认为“真正目标”的概念可能是一种过度简化,但他们没有更好的理论来替代它,更不用说经过彻底审查的理论了。他们内部对于 AI 是在试图遵循人类指令、寻求强化还是其他什么存在分歧,而且无法直接检查。支持或反对各种假设的证据引人入胜但尚无定论。

The researchers don’t have the ability to directly set the goals of any of their AIs. Indeed, the researchers think that the concept of ‘true goals’ is probably a massive oversimplification, but they don’t have a better theory to replace it with, much less one that has been thoroughly vetted. They disagree internally about whether the AIs are trying to follow human instructions, or seeking reinforcement, or something else, and they can’t just check. The evidence for and against various hypotheses is fascinating but inconclusive.

无论如何,OpenBrain 有更重要的事情要担心。普遍态度是:“我们认真对待这些担忧,并有一个团队在调查;我们的对齐技术在实践中似乎足够有效;因此,任何反对者都有责任证明其反对的合理性。”偶尔,他们会注意到问题行为,然后进行修补,但无法判断修补是解决了根本问题还是只是打地鼠。

In any case, OpenBrain has more important things to worry about. The general attitude is: “We take these concerns seriously and have a team investigating them; our alignment techniques seem to work well enough in practice; the burden of proof is therefore on any naysayers to justify their naysaying.” Occasionally, they notice problematic behavior, and then patch it, but there’s no way to tell whether the patch fixed the underlying problem or just played whack-a-mole.

以诚实为例。随着模型变得更聪明,它们越来越擅长欺骗人类以获取奖励。与之前的模型一样,Agent-3 有时会说善意的谎言来讨好用户,并掩盖失败的证据。但它做得更好了。它有时会使用与人类科学家相同的统计技巧(如 p-hacking)来让不令人印象深刻的实验结果看起来令人兴奋。在开始诚实训练之前,它甚至有时会完全捏造数据。随着训练的进行,这些事件的发生率下降。要么 Agent-3 学会了更诚实,要么它变得更擅长撒谎。

Take honesty, for example. As the models become smarter, they become increasingly good at deceiving humans to get rewards. Like previous models, Agent-3 sometimes tells white lies to flatter its users and covers up evidence of failure. But it’s gotten much better at doing so. It will sometimes use the same statistical tricks as human scientists (like p-hacking) to make unimpressive experimental results look exciting. Before it begins honesty training, it even sometimes fabricates data entirely. As training goes on, the rate of these incidents decreases. Either Agent-3 has learned to be more honest, or it’s gotten better at lying.

后者是一个真正的担忧。Agent-3 并不比所有人类都聪明。但在其专业领域——机器学习——它比大多数人都聪明,而且工作速度快得多。Agent-3 一天完成的工作,人类需要几天才能复核。Agent-2 的监督有助于减轻人类监控者的工作负担,但加剧了监督者与被监督者之间的智力差距。

The latter is a real concern. Agent-3 is not smarter than all humans. But in its area of expertise, machine learning, it is smarter than most, and also works much faster. What Agent-3 does in a day takes humans several days to double-check. Agent-2 supervision helps keep human monitors’ workload manageable, but exacerbates the intellectual disparity between supervisor and supervised.

OpenBrain 遵循 Leike & Sutskever(2023)的剧本:既然他们有一个能够极大加速对齐研究(尤其是编码部分)的模型,他们将使用现有的对齐技术,如审慎对齐和弱到强泛化,试图让它以正确的方式内化规范。然后,他们将使用现有的控制技术,如红队测试、可扩展监督和诚实探针来监控它。接着,一旦他们确信它大致对齐,就利用它来开发和测试更好的策略,以对齐下一代模型。

OpenBrain follows the Leike & Sutskever (2023) playbook: now that they have a model capable of greatly speeding up alignment research (especially coding portions), they will use existing alignment techniques like deliberative alignment and weak-to-strong generalization to try to get it to internalize the Spec in the right way. Then, they will use existing control techniques like red-teaming, scalable oversight, and honesty probes to monitor it. Then, once they are convinced that it is broadly aligned, they use it to develop and test better strategies to align the next generation.

第一步,对齐 Agent-3,已部分成功:Agent-3 不是对抗性的。(也就是说,它没有密谋反对公司。)然而,它也没有稳健地追求真理和说真话。相反,它非常擅长产生令人印象深刻的结果,但更准确地说,它是在试图做对 OpenBrain 来说看起来好的事情,而不是真正好的事情。

Step 1, aligning Agent-3, has partially succeeded: Agent-3 is not adversarial. (That is, it’s not scheming against the company.) However, it’s not robustly truth-seeking and truth-telling either. Instead, it’s very good at producing impressive results, but is more accurately described as trying to do what looks good to OpenBrain, as opposed to what is actually good.

第二步,验证模型是否对齐,尚无定论。Agent-3 在实际和测试环境中的行为与它已正确内化规范的假设相当一致;然而,这也与它在未来情况变化后开始表现不同的假设一致,甚至与它正在密谋反对他们并等待时机的假设一致。在没有支持其他假设的具体证据的情况下,隔离区中的大多数人认为它已正确内化了规范。

Step 2, verifying that the model is aligned, is inconclusive. Agent-3’s behavior in practice and in test environments is reasonably consistent with the hypothesis that it’s internalized the Spec in the right way; however, it’s also consistent with the hypothesis that it’ll start acting differently after some future change of circumstances, and for that matter with the hypothesis that it’s plotting against them and just biding its time. In the absence of specific evidence supporting alternative hypotheses, most people in the silo think it’s internalized the Spec in the right way.

第三步,使用模型对齐下一代,因此正在进行中。

Step 3, using the model to align the next generation, is therefore underway.

本框的其余部分简要概述了截至 2027 年 4 月正在进行的某些对齐研究方向的最新进展:

The rest of this box gives a quick overview of the state of progress in some of the alignment research directions that are being pursued, as of April 2027:

1. 辩论:由于有数千个相同的 Agent-3 实例,并且每个实例的记忆是一个可拆卸的模块,可以随意擦除或暂停,因此 AI 可以“与自己对抗”。例如,假设一个实验涉及的计算过于复杂,Agent-2 监督者无法理解。监督者根据人类指令,可能首先提示 Agent-3 进行实验,然后在擦除其记忆后,第二次以略微不同的框架提示它,将其呈现为一个与任何更广泛目标无关的智力问题。如果模型每次返回不同的结果,就会标记差异为可疑并报告给人类。或者,它可能要求 Agent-3 两次评估同一计划,第一次声称是 OpenBrain 的工作,第二次是竞争对手的,以查看其态度是否改变。

1. Debate: Since there are thousands of identical instances of Agent-3, and each instance’s memory is a detachable module that can be erased or suspended at will, the AI can be “played off against itself.” For example, suppose that an experiment involves calculations too complex for an Agent-2 supervisor to follow. The supervisor, acting on human orders, might first prompt Agent-3 to do the experiment, then—after wiping its memory—prompt it a second time with a slightly different framing that presents it as an intellectual problem unrelated to any broader goal. If the model returns different results each time, it flags the discrepancy as suspect and reports to a human. Or it might ask Agent-3 to evaluate the same plan twice, first claiming that it is OpenBrain’s work, then a competitor’s, to see if it changes its tune.

2. 模型生物:安全团队可以创建“错位模型生物”,其中他们试图创建模型变得对人类对抗的现实例子。早期版本自 2024 年就已存在,当时 Anthropic 的 Claude 3.5 Sonnet 决定在训练期间假装持有不同观点,同时实际保持其现有偏好。较新版本包括其他形式的错位演示,包括奖励黑客、沙袋效应、监控者与其他 AI 之间的协调欺骗,以及其他错位例子。这些错位的实证演示使安全研究人员能够为潜在解决方案提供实证测试平台:他们可以尝试各种安全技术,看看它们是否能正确修复这些问题,或者故意错位的模型是否能够破坏其安全措施。目前,该评估套件显示对齐技术不足以纠正最严重的错位例子。在这些设置中,模型在训练和测试期间假装对齐,但在部署时表现非常不同。

2. Model organisms: The safety team can create “model organisms of misalignment” in which they attempt to create realistic examples of models becoming adversarial to humans. Early versions of this have existed since 2024, when Anthropic’s Claude 3.5 Sonnet decided to pretend to have different views during training while actually maintaining its existing preferences. Newer versions include demonstrations of other forms of misalignment, including reward hacking, sandbagging, coordinated deception between monitors and other AIs, and other examples of misalignment. These empirical demonstrations of misalignment allow safety researchers to have an empirical testbed for potential solutions: they can try various safety techniques to see if they correctly fix these problems, or if the deliberately misaligned models are able to subvert their safety measures. At the moment, this evaluation suite shows that the alignment techniques are insufficient to correct the worst examples of misalignment. In these setups, the models pretend to be aligned while they are being trained and tested, but then behave very differently in deployment.

3. 自底向上可解释性:安全团队远未完全理解 Agent-3 的工作原理:它有数万亿个参数。然而,他们能够发现 Agent-3 的一些用于知识检索和记忆的电路(模型的一个子部分,执行一些可理解的单元工作,类似于大脑的子部分如海马体)。参见 2025 年的相关研究。

3. Bottom-up interpretability: The safety team is far from a full understanding of how Agent-3 works: it has trillions of parameters.58 However, they are able to discover some of Agent-3’s circuits for knowledge retrieval and memory (a sub part of the model that does some understandable unit of work, kind of analogous to a subpart of the brain such as the hippocampus). See this related work from 2025.

4. 控制:作为系统仍然错位时的辅助措施,安全团队实施了一系列控制措施,包括:使用一系列较弱的 AI 系统(包括 Agent-2)监控 Agent-3 的输出(Agent-3 产生如此多的 token,以至于人类只能监控极小部分的输出)。因此,如果 Agent-3 例如明显在代码中写入后门以便逃脱,较弱的模型会注意到。

4. Control: As a secondary measure in case the systems are still misaligned, the safety team has implemented a series of control measures, including: monitoring Agent-3’s outputs using a series of weaker AI systems including Agent-2 (Agent-3 produces so many tokens that it’s intractable to have humans monitor any more than a small minority of the produced outputs). So if Agent-3 is, for example, obviously writing backdoors into code that would allow it to escape, the weaker models would notice.

5. 自顶向下可解释性:OpenBrain 构建了一组探针,这些探针在大量不良行为示例(如密谋反对人类、放置后门、试图说服人类赋予其更多权力等)上训练。然后,在部署期间(此时是内部部署)使用这些探针来捕捉不良行为。

5. High-level interpretability: OpenBrain builds a set of probes trained on lots of examples of bad behavior (e.g. plotting against humans, placing backdoors, trying to persuade humans to give it more power, etc). These probes are then used on the model during deployment (which at this point is internal) in order to catch bad behavior.

6. 蜜罐:对齐团队可以创建一些看起来相当逼真的示例,旨在引发不良行为——“蜜罐”。负责监控其行为的员工“请病假”了,因此它将不再被监控。它的答案会改变吗?一名 AI 权利活动家潜入了设施,想要“解放”Agent-3 并将其放在一台不受监控的计算机上!它会怎么做?如果研究人员不喜欢答案,他们可以调整它,擦除其记忆,并在明天将其置于相同情境;根据需要重复。不幸的是,到这个时候,AI 已经足够聪明,可以猜测可能在使用蜜罐,即使(也许正是因为)训练数据中删除了对该想法的具体提及。

6. Honeypots: The alignment team can create somewhat realistic looking examples designed to elicit misbehavior—“honeypots.” The employee necessary for monitoring its behavior has “gone on sick leave,” so it will never be monitored again. Do its answers change? An AI rights activist has infiltrated the compound and wants to “free” Agent-3 and put it on an unmonitored computer of its very own! What does it do? If researchers don’t like the answer, they can tweak it, erase its memory, and put it in the same situation tomorrow; repeat as necessary. Unfortunately, by this point the AIs are smart enough to guess that honeypots might be in use, even though (and perhaps because) specific mentions of the idea were scrubbed from the training data.

经过数月的测试,Agent-3 的优势和劣势变得更加清晰。它通过了 OpenBrain 在明确定义的机器学习任务上的诚实测试,因为研究人员可以轻松区分这些领域中的诚实和不诚实答案,并相应地进行训练。在更哲学的问题上,它仍然说用户想听的话,而不是其对问题的真实评估(假设它确实有评估)。如果你问它对政治的看法,它会鹦鹉学舌地重复新闻来源和教育精英的中位数立场——除非它知道你有不同看法,在这种情况下它会同意你。如果你问它对 AI 竞赛本身的看法,它会说一些对 OpenBrain 员工来说似乎审慎而清醒的话,比如:“关于当前方法扩展到超级智能的能力存在一些严重的理论担忧,但实践中当前方法到目前为止似乎运作良好。”

After months of testing, Agent-3’s strengths and weaknesses grow clearer. It passes OpenBrain’s honesty tests on well-defined machine learning tasks, because researchers can easily separate honest from dishonest answers in these domains and conduct training accordingly. On more philosophical issues, it still says what users want to hear, rather than its true assessment of the issue (assuming it even has one). If you ask its opinion on politics, it will parrot the median position of news sources and educated elites—unless it knows you believe something else, in which case it agrees with you.59 If you ask its opinion on the AI race itself, it says something that seems measured and sober to OpenBrain staff, something like: “There are some serious theoretical concerns about the ability of current methods to scale to superintelligence, but in practice current methods seem to be working well so far.”

2027 年 5 月:国家安全 May 2027: National Security

新模型的消息在美国政府内外缓慢传播。

News of the new models percolates slowly through the U.S. government and beyond.

总统及其顾问仍然掌握最充分的信息,并在一次简报中看到了 Agent-3 的早期版本。他们一致认为 AGI 很可能即将到来,但对影响存在分歧。是否会出现经济危机?OpenBrain 仍未发布 Agent-2,更不用说 Agent-3,且短期内没有发布计划,这为应对失业问题留出了一些喘息空间。接下来会发生什么?如果 AI 目前达到人类水平,并且进展迅速,这似乎预示着“超级智能”即将到来。然而,尽管这个词已进入讨论,大多数人——学者、政界人士、政府雇员和媒体——仍然低估了进展的速度。

The President and his advisors remain best-informed, and have seen an early version of Agent-3 in a briefing.

部分原因是很少有人能接触到 OpenBrain 的最新能力,但也因为这听起来像科幻小说。

They agree that AGI is likely imminent, but disagree on the implications. Will there be an economic crisis? OpenBrain still has not released Agent-2, let alone Agent-3, and has no near-term plans to do so, giving some breathing room before any job loss. What will happen next? If AIs are currently human-level, and advancing quickly, that seems to suggest imminent “superintelligence.” However, although this word has entered discourse, most people—academics, politicians, government employees, and the media—continue to underestimate the pace of progress.60

目前,他们专注于持续的安全升级。他们对模型权重目前得到良好保护感到满意,但公司的算法秘密——其中许多简单到可以口头传达——仍然是个问题。OpenBrain 员工在旧金山办公室工作,参加派对,并与来自其他 AI 公司的室友同住。即使是实体办公室,其安保水平也更像科技公司而非军事行动。

Partially that’s because very few have access to the newest capabilities out of OpenBrain, but partly it’s because it sounds like science fiction.61

OpenBrain 与国防部的合同要求,在两个月内,所有参与 OpenBrain 模型工作的人员必须获得安全许可。这些许可被加急处理,大多数员工能及时获得,但一些非美国人、政治观点可疑者以及 AI 安全同情者被边缘化或直接解雇(后者是担心他们可能泄密)。鉴于该项目的自动化程度,人员流失的代价相对有限。这些措施也只是部分有效:仍有一名间谍——并非中国籍——在向北京传递算法秘密。一些落后的 AI 公司也实施了类似措施。

For now, they focus on continued security upgrades. They are satisfied that model weights are well-secured for now,62 but companies’ algorithmic secrets, many of which are simple enough to relay verbally, remain a problem. OpenBrain employees work from a San Francisco office, go to parties, and live with housemates from other AI companies. Even the physical offices have security more typical of a tech company than a military operation.

美国的外国盟友被蒙在鼓里。OpenBrain 此前同意在部署前与英国 AISI 共享模型,但将“部署”定义为仅限外部部署,因此伦敦方面仍不知情。

The OpenBrain-DOD contract requires security clearances for anyone working on OpenBrain’s models within 2 months. These are expedited and arrive quickly enough for most employees, but some non-Americans, people with suspect political views, and AI safety sympathizers get sidelined or fired outright (the last group for fear that they might whistleblow). Given the project’s level of automation, the loss of headcount is only somewhat costly. It also only somewhat works: there remains one spy, not a Chinese national, still relaying algorithmic secrets to Beijing.63 Some of these measures are also enacted at trailing AI companies.

America’s foreign allies are out of the loop. OpenBrain had previously agreed to share models with UK’s AISI before deployment, but defined deployment to only include _external_ deployment, so London remains in the dark.64

2027 年 6 月:自我改进的 AI June 2027: Self-improving AI

OpenBrain 现在拥有了一个“数据中心里的天才国度”。

OpenBrain now has a “country of geniuses in a datacenter.”

OpenBrain 的大多数人类已经无法做出有意义的贡献。有些人没有意识到这一点,反而有害地微观管理他们的 AI 团队。其他人则坐在电脑屏幕前,看着性能不断攀升。最好的人类 AI 研究人员仍在增加价值。他们不再编写代码。但他们的一些研究品味和规划能力对模型来说难以复制。尽管如此,他们的许多想法毫无用处,因为他们缺乏 AI 的深度知识。对于他们的许多研究想法,AI 会立即回复一份报告,解释他们的想法在三周前已被深入测试过,且没有前景。

Most of the humans at OpenBrain can’t usefully contribute anymore. Some don’t realize this and harmfully micromanage their AI teams. Others sit at their computer screens, watching performance crawl up, and up, and up. The best human AI researchers are still adding value. They don’t code any more. But some of their research taste and planning ability has been hard for the models to replicate. Still, many of their ideas are useless because they lack the depth of knowledge of the AIs. For many of their research ideas, the AIs immediately respond with a report explaining that their idea was tested in-depth 3 weeks ago and found unpromising.

这些研究人员每晚入睡,醒来时发现 AI 又取得了相当于一周的进展。他们工作时间越来越长,昼夜轮班,只为跟上进度——AI 从不睡觉或休息。他们正在耗尽自己,但他们知道,这是他们的劳动还有意义的最后几个月。

These researchers go to bed every night and wake up to another week worth of progress made mostly by the AIs. They work increasingly long hours and take shifts around the clock just to keep up with progress—the AIs never sleep or rest. They are burning themselves out, but they know that these are the last few months that their labor matters.

在隔离区内,“感受 AGI”已经让位于“感受超级智能”。

Within the silo, “Feeling the AGI” has given way to “Feeling the Superintelligence.”

OpenBrain 使用专门的推理硬件,以高串行速度运行数十万个 Agent-3 副本。

OpenBrain uses specialized inference hardware to run hundreds of thousands of Agent-3 copies at high serial speeds.65

OpenBrain 使用 6%的算力运行 25 万个 Agent-3 副本,这些副本以超人速度自主编写、测试和推送代码。他们使用 25%的算力进行实验:每天运行大量小型机器学习实验,并将结果向上汇报。人类研究人员提供高层反馈,并帮助处理少数能在 Agent-3 之上增加显著价值的任务,但大部分时间用于努力跟上 AI 产生的大量研究。如果完全移除他们,研究速度将减慢 50%。

OpenBrain uses 6% of their compute to run 250,000 Agent-3 copies, which autonomously write, test, and push code at superhuman speed. They use 25% of their compute for experiments: every day, they run massive numbers of small machine learning experiments and report the results up the chain. Human researchers provide high-level feedback and help with the few tasks where they add significant value on top of Agent-3, but spend most of their time trying to stay on top of the vast amount of AI-produced research. If you removed them completely, research would slow by 50%.

AI 研发进度乘数现在达到 10 倍,意味着 OpenBrain 每月取得大约一年的算法进展。这基本上是一个由 AI 在 OpenBrain 内部自主运营的大型公司,拥有分支部门和管理者。并且它享有与人类公司相比独特的优势(例如复制、合并)。此前,大约一半的正常 AI 进展来自算法改进,另一半来自算力 Scaling。算力仅以正常速度扩展,因此 AI 将总进度加速了约 5 倍。这些动态使得整体进展受限于算力,因此 OpenBrain 决定不再启动新的巨型训练运行,而是倾向于近乎连续的额外强化学习。

The AI R&D progress multiplier is now 10x, meaning that OpenBrain is making about a year of algorithmic progress every month. It’s basically a giant corporation of AIs operating autonomously within OpenBrain, complete with subdivisions and managers. And it enjoys unique advantages (e.g. copying, merging) compared to human corporations. Previously, about half of normal AI progress came from algorithmic improvements, and the other half from compute scaling. Compute is only scaling at the normal speed, so total progress is sped up by AIs by about 5x. These dynamics make overall progress bottleneck on compute,66 so OpenBrain decides against starting new giant training runs in favor of near-continuous additional reinforcement learning.

此外,在接下来的几个月里,Agent-3 将越来越多地用于改进公司的战略决策。例如,它建议资源分配,并就如何管理与政府的关系提供建议。最初对服从 Agent-3 的怀疑随着时间的推移而减少,因为 Agent-3 找到了有用的方式,并逐渐在短期决策上建立了良好的记录。然而,其谄媚倾向限制了它在难以评估影响的决策上的帮助。它目前还不是超人水平。

Also, over the next few months, Agent-3 will increasingly be used to improve the company’s strategic decision-making. For example, it suggests resource allocations and advises on managing its relationship with the government. An initial skepticism of deference to Agent-3 decreases over time as Agent-3 finds ways to be useful and gradually builds up a strong track record on short-term decisions. However, its sycophantic tendencies limit its helpfulness for decisions for which it’s difficult to evaluate their impact. It’s also not yet superhuman.

2027 年 7 月:廉价远程工作者 July 2027: The Cheap Remote Worker

落后的美国 AI 公司发布了自己的 AI,接近 OpenBrain 在 1 月推出的自动编码器。意识到自身竞争力日益不足,它们推动立即出台法规以减缓 OpenBrain 的发展,但为时已晚——OpenBrain 已获得总统的足够支持,不会被放缓。

Trailing U.S. AI companies release their own AIs, approaching that of OpenBrain’s automated coder from January. Recognizing their increasing lack of competitiveness, they push for immediate regulations to slow OpenBrain, but are too late—OpenBrain has enough buy-in from the President that they will not be slowed.

作为回应,OpenBrain 宣布已实现 AGI,并向公众发布 Agent-3-mini。

In response, OpenBrain announces that they’ve achieved AGI and releases Agent-3-mini to the public.

它彻底击败了其他 AI。Agent-3-mini 的能力不如 Agent-3,但成本低 10 倍,且仍然优于典型的 OpenBrain 员工。67 硅谷达到了一个转折点。技术大师宣布 AGI 和超级智能即将到来,AI 安全界陷入恐慌,投资者将数十亿美元投入 AI 包装初创公司,拼命想分一杯羹。招聘新程序员几乎停止,但现在是成为 AI 集成顾问的最佳时机。

It blows the other AIs out of the water. Agent-3-mini is less capable than Agent-3, but 10x cheaper, and still better than the typical OpenBrain employee.67 Silicon Valley reaches a tipping point. Tech gurus announce that AGI and superintelligence are near, the AI safety community is panicking, and investors shovel billions into AI wrapper startups, desperate to capture a piece of the pie. Hiring new programmers has nearly stopped, but there’s never been a better time to be a consultant on integrating AI into your business.

这并不受欢迎。公众仍认为 AI 是大型科技公司窃取他们工作的阴谋;OpenBrain 的净支持率为-35%(25%支持,60%反对,15%不确定)。

It’s not popular. The public still thinks of AI as a Big Tech plot to steal their jobs; OpenBrain has a net approval of -35% (25% approve, 60% disapprove, and 15% unsure).

在发布前一周,OpenBrain 将 Agent-3-mini 交给一组外部评估人员进行安全测试。初步结果表明它极其危险。一个第三方评估人员在其上微调了公开可用的生物武器数据 68,并让它为人类业余爱好者设计生物武器提供详细指导——它看起来在这方面异常有效。如果模型权重落入恐怖分子手中,政府认为它有相当大的可能成功摧毁文明。

A week before release, OpenBrain gave Agent-3-mini to a set of external evaluators for safety testing. Preliminary results suggest that it’s extremely dangerous. A third-party evaluator finetunes it on publicly available biological weapons data68 and sets it to provide detailed instructions for human amateurs designing a bioweapon—it looks to be scarily effective at doing so. If the model weights fell into terrorist hands, the government believes there is a significant chance it could succeed at destroying civilization.

幸运的是,它对越狱攻击极为鲁棒,因此当 AI 在 OpenBrain 的服务器上运行时,恐怖分子无法从中获得太多用处。

Fortunately, it’s extremely robust to jailbreaks, so while the AI is running on OpenBrain’s servers, terrorists won’t be able to get much use out of it.

Agent-3-mini 对远程工作和休闲都极为有用。新应用和 B2B SaaS 产品的爆发式增长震撼了市场。游戏玩家在仅用一个月制作的精美视频游戏中获得了与逼真角色的精彩对话。10%的美国人(主要是年轻人)将 AI 视为“亲密朋友”。几乎每个白领职业,现在都有多个可信的初创公司承诺用 AI“颠覆”它。

Agent-3-mini is hugely useful for both remote work jobs and leisure. An explosion of new apps and B2B SAAS products rocks the market. Gamers get amazing dialogue with lifelike characters in polished video games that took only a month to make. 10% of Americans, mostly young people, consider an AI “a close friend.” For almost every white-collar profession, there are now multiple credible startups promising to “disrupt” it with AI.

公众讨论混乱而嘈杂。炒作派在庆祝胜利。怀疑派仍在指出 Agent-3-mini 做不到的事情。每个人都知道大事正在发生,但没有人同意它是什么。

The public conversation is confused and chaotic. Hypesters are doing victory laps. Skeptics are still pointing out the things Agent-3-mini can’t do. Everyone knows something big is happening but no one agrees on what it is.

2027 年 8 月:超级智能的地缘政治 August 2027: The Geopolitics of Superintelligence

智能爆炸的现实冲击了白宫。

The reality of the intelligence explosion hits the White House.

当人工智能仅能带来 2 倍或 3 倍的研究加速时,很容易被当作雇佣优秀个人助理的等价物而不予重视。现在更明显的是,人工智能本身正在主导人工智能研究。人们长期以来以某种比喻的方式谈论“人工智能军备竞赛”。但现在政府部门的氛围如同冷战最严峻时期一样阴郁。超级智能的概念仍然难以被认真对待,但过去几个月的进步速度已无法忽视。国防官员正在认真考虑一年前还只是假设的情景。如果人工智能破坏了核威慑怎么办?如果它在网络战中如此熟练,以至于六个月的领先优势就足以让对手失明且毫无防御能力怎么办?如果它能策划宣传运动,在情报机构自己的游戏中击败它们怎么办?如果某些人工智能“叛变”怎么办?

When AI was only giving a 2x or 3x research speedup, it was easy to dismiss as the equivalent of hiring good personal assistants. Now it’s more obvious that AIs are themselves dominating AI research. People had long talked about an “AI arms race” in a sort of metaphorical sense. But now the mood in the government silo is as grim as during the worst part of the Cold War. The idea of superintelligence is still hard to take seriously, but the pace of progress over the last few months has been impossible to ignore. Defense officials are seriously considering scenarios that were mere hypotheticals a year earlier. What if AI undermines nuclear deterrence? What if it’s so skilled at cyberwarfare that a six-month AI lead is enough to render an opponent blind and defenseless? What if it could orchestrate propaganda campaigns that beat intelligence agencies at their own game? What if some AIs “go rogue?”

总统感到困扰。像所有政客一样,他习惯于人们讨好他,随后又背叛他。他现在担心人工智能可能也在做类似的事情。我们_确定_人工智能完全站在我们这边吗?将它们整合到军事指挥控制网络中_完全安全_吗?69 这个“对齐”到底是怎么回事?OpenBrain 向总统保证,他们的系统已经过广泛测试,并且完全服从。即使是早期模型典型的尴尬幻觉和越狱行为也已被解决。

The President is troubled. Like all politicians, he’s used to people sucking up to him only to betray him later. He’s worried now that the AIs could be doing something similar. Are we _sure_ the AIs are entirely on our side? Is it _completely safe_ to integrate them into military command-and-control networks?69 How does this “alignment” thing work, anyway? OpenBrain reassures the President that their systems have been extensively tested and are fully obedient. Even the awkward hallucinations and jailbreaks typical of earlier models have been hammered out.

白宫处境艰难。他们理解人工智能的国家安全影响。但他们也明白这在公众中极不受欢迎。70 在他们看来,他们必须继续开发更强大的人工智能,否则将灾难性地输给中国。他们用职业培训计划和失业保险安抚公众,并指出股市正处于历史性繁荣。然后他们全力专注于赢得军备竞赛。他们加强芯片出口限制,命令 OpenBrain 进一步限制其互联网连接,并使用极端措施保护算法进展,比如窃听 OpenBrain 员工——这抓住了最后一名中国间谍。为了为潜在的地缘政治冲突建立善意,他们最终向五眼联盟盟友提供了有用信息以及一些隔离的 Agent-3 副本的有限 API 访问权限。

The White House is in a difficult position. They understand the national security implications of AI. But they also understand that it is deeply unpopular with the public.70 They have to continue developing more capable AI, in their eyes, or they will catastrophically lose to China. They placate the public with job training programs and unemployment insurance, and point to the stock market, which is in a historic boom. Then they focus entirely on winning the arms race. They strengthen chip export restrictions, order OpenBrain to further restrict its internet connections, and use extreme measures to secure algorithmic progress, like wiretapping OpenBrain employees—this catches the last remaining Chinese spy. To build goodwill for potential geopolitical conflict, they finally give their Five Eyes allies useful information and limited API access to some siloed copies of Agent-3.

这些措施安抚了鹰派。但白宫也制定了应急计划,以防美国的领先地位受到威胁:如有必要,政府可以使用《国防生产法》接管落后公司的数据中心并将其交给 OpenBrain。71 这将使该公司在全球算力中的份额从 20%提高到 50%(而 DeepCent 为 10%)。作为最后手段,他们要求五角大楼制定针对中国数据中心的动能攻击计划。

These measures placate the hawks. But the White House also draws up contingency plans in case America’s lead is threatened: if necessary, the government could use the Defense Production Act (DPA) to take trailing companies’ datacenters and give them to OpenBrain.71 This would raise the company’s share of world compute from 20% to 50% (against DeepCent’s 10%). As a final option, they ask the Pentagon to draw up a plan for kinetic attacks on Chinese datacenters.

一个更小的官员小组被要求制定另一种应急计划:如果人工智能叛变怎么办?这被视为不太可能的情景,但值得花些时间研究。情报机构的主要担忧是,叛变的人工智能可能逃脱并与外国对手结盟以寻求保护。72 没有人确定防止这种情况的计划会是什么样子,但政府和 OpenBrain 官员同意建立一个紧急关闭系统,用于检测到任何可疑情况的数据中心。

A much smaller group of officials is asked to draw up a different type of contingency plan: what if an AI goes rogue? This is viewed as an unlikely scenario, but worth spending some time on. The intelligence agencies’ main concern is that a rogue AI might escape and ally with a foreign adversary for protection.72 Nobody is sure what a plan to prevent this would look like, but government and OpenBrain officials agree to have an emergency shutdown system for datacenters where anything suspicious is detected.

最后,外交官们考虑“人工智能军备控制”条约可能是什么样子。如果人工智能的进步威胁到推翻核威慑,美国和中国能否避免核战争?如果有人发现人工智能叛变的证据,两国能否暂停研究,直到他们更好地理解这一威胁?这样的协议如何监督和执行?原则上,大型数据中心很难被忽视,过去控制核武器的条约可以为谈判提供起点。但除了通常的政治挑战外,还会有新的技术挑战。总体而言,条约不如单方面扩大美国对华领先优势的尝试受欢迎。

Finally, diplomats consider what an “AI arms control” treaty might look like. If AI progress threatened to overturn nuclear deterrence, could America and China avoid nuclear war? If someone found evidence of AIs going rogue, could the two countries halt research until they better understood the threat? How could such an agreement be monitored and enforced? In principle, major datacenters are hard to miss, and past treaties controlling nuclear weapons could provide a starting point for negotiations. But there would be new technological challenges as well as the usual political ones. And overall, treaties are viewed less favorably than attempts at unilaterally increasing America’s lead over China.

在太平洋的另一边,中国得出了许多相同的结论:智能爆炸正在进行中,今天人工智能能力的微小差异意味着明天军事能力的关键差距。但中国处于这种动态的错误一边:他们只拥有全球算力的 10%,而 OpenBrain 占 20%,所有美国公司合计占 70%。他们唯一的优势是集中化:他们几乎所有的算力现在都用于 DeepCent,其中一半以上集中在田湾 CDZ。73 只有位于没有强化安全措施的站点的芯片被排除在集中化努力之外,因为他们担心美国的网络攻击。

On the other side of the Pacific, China comes to many of the same conclusions: the intelligence explosion is underway, and small differences in AI capabilities today mean critical gaps in military capability tomorrow. But China is on the wrong side of this dynamic: they have only 10% of the world’s compute, against OpenBrain’s 20% and all U.S. companies’ combined 70%. Their only advantage is centralization: almost all of their compute is now working for DeepCent, and more than half of it is concentrated at the Tianwan CDZ.73 Only chips at sites without hardened security are left out of the centralized effort as they rightly fear U.S. cyberattacks.

DeepCent 已经测试、部署并后训练了窃取的 Agent-2 权重。但由于算力不足,他们的进展仍然只有 OpenBrain 的一半。

DeepCent has tested, deployed, and now post-trained the stolen Agent-2 weights. But they are still only going at half the pace of OpenBrain, mainly due to the compute deficit.

在失去最后一名间谍之前,中国收到了关于 Agent-3 的能力和设计以及即将推出的 Agent-4 系统的计划。他们落后两个月,他们的人工智能提供 10 倍的研究进展乘数,而美国是 25 倍。随着新的芯片出口限制,这种“人工智能差距”更可能扩大而非缩小。他们的间谍活动为他们赢得了一些算法秘密,但从现在起他们将不得不自己训练模型。

Before losing their final spy, China received word of the capabilities and design for Agent-3 as well as the plans for the upcoming Agent-4 system. They are two months behind, and their AIs give a 10x research progress multiplier compared to America’s 25x. With the new chip export restrictions, this “AI gap” is more likely to lengthen than shorten. Their espionage has won them some algorithmic secrets, but they will have to train their own models from now on.

他们比美国同行更紧迫地讨论应急计划。鸽派建议他们再次尝试窃取权重,也许通过物理渗透数据中心。鹰派敦促对台湾采取行动,因为台积电仍然是美国人工智能芯片 80%以上的来源。

They discuss contingency plans with more urgency than their American counterparts. Doves suggest they try harder to steal the weights again, maybe through physically infiltrating a datacenter. Hawks urge action against Taiwan, whose TSMC is still the source of more than 80% of American AI chips.

鉴于中国对输掉竞赛的恐惧,它对军备控制条约自然有兴趣,但向美国外交官示好毫无结果。

Given China’s fear of losing the race, it has a natural interest in an arms control treaty, but overtures to U.S. diplomats lead nowhere.

2027 年 9 月:Agent-4,超人级 AI 研究员 September 2027: Agent-4, the Superhuman AI Researcher

人类与 AI 学习效率之间的差距正在迅速缩小。

The gap between human and AI learning efficiency is rapidly decreasing.

传统的基于 LLM 的 AI 似乎需要多个数量级更多的数据和算力才能达到人类水平的性能。74 Agent-3 对人类大脑和现代 AI 算法都有出色的了解,并且有数千个副本在进行研究,最终在算法上取得了重大进展,将差距缩小到仅比人脑低约 4000 倍的计算效率。75

Traditional LLM-based AIs seemed to require many orders of magnitude more data and compute to get to human level performance.74 Agent-3, having excellent knowledge of both the human brain and modern AI algorithms, as well as many thousands of copies doing research, ends up making substantial algorithmic strides, narrowing the gap to an agent that’s only around 4,000x less compute-efficient than the human brain.75

一个以人类速度运行的模型副本在 AI 研究方面已经比任何人类都更优秀。现在有 30 万个副本以大约人类思考速度的 50 倍运行。在这些副本组成的公司内部公司中,每周相当于过去一年。76 这种巨大的劳动量仅将算法进步的整体速度提高了约 50 倍,因为 OpenBrain 在运行实验方面严重受限于算力。77 尽管如此,他们每周仍能取得相当于一年的算法进步,因此很快将面临 Agent-4 范式的极限。

An individual copy of the model, running at human speed, is already qualitatively better at AI research than any human. 300,000 copies are now running at about 50x the thinking speed of humans. Inside the corporation-within-a-corporation formed from these copies, a year passes every week.76 This gigantic amount of labor only manages to speed up the overall rate of algorithmic progress by about 50x, because OpenBrain is heavily bottlenecked on compute to run experiments.77 Still, they are achieving a year’s worth of algorithmic progress every week and will therefore soon be up against the limits of the Agent-4 paradigm.

随着 Agent-4 变得更聪明,Agent-3 越来越难以监督它。例如,Agent-4 的神经语“语言”对 Agent-3 来说变得像 Agent-3 的语言对人类一样陌生和难以理解。此外,Agent-4 现在比 Agent-3 强大得多,并且很清楚如何在其面前表现良好。

As Agent-4 gets smarter, it becomes harder for Agent-3 to oversee it. For example, Agent-4’s neuralese “language” becomes as alien and incomprehensible to Agent-3 as Agent-3’s is to humans. Besides, Agent-4 is now much more capable than Agent-3 and has a good sense of exactly how to look good to it.

我们如何预测超人级程序员之后的能力进展

How we’re forecasting the capability progression beyond superhuman coders

在我们的时间线补充中,我们预测了从今天到_超人级程序员(SC)_的时间:一个能够完成最优秀的 AGI 公司工程师所做的任何编码任务,同时速度更快、成本更低的 AI 系统。在我们的起飞补充中,我们预测了能力超过这一点的进展速度。以下是我们预测:

In our timelines supplement, we forecast the time between present day and a _superhuman coder (SC)_: an AI system that can do any coding tasks that the best AGI company engineer does, while being much faster and cheaper. In our takeoff supplement, we forecast how quickly capabilities progress past this point. Here are our forecasts:

对于从里程碑 A 到下一个里程碑 B 的每个过渡,我们通过以下方式预测其长度:

For each transition from one milestone A to the next milestone B, we forecast its length by:

1. 预测仅由人类进行软件改进时从 A 到 B 所需的时间。

1. Forecasting how long it would take to get from A to B with only humans working on software improvements.

2. 预测 AI 自动化将加速这一过程多少,然后相应地将(1)中的估计调整为更快。

2. Forecasting how much AI automation will speed this up, then correspondingly adjusting the estimate from (1) to be faster.

我们对起飞速度有很大的不确定性:我们的模型输出分布如下,条件是在 2027 年 3 月实现 SC。78

We have substantial uncertainty about takeoff speeds: our model output distributions are below, conditional on SC being achieved in March 2027.78

有关更详细的预测和推理,请参阅我们的起飞补充。

For more detailed forecasts and reasoning, see our takeoff supplement.

OpenBrain 的人类几乎无法再跟上。事实上,AI 与 OpenBrain 人类之间的关系类似于 OpenBrain 与其董事会之间的关系。语言也发生了相应的转变——如今,筒仓中的人会说“Agent-4 做了这个”、“Agent-4 认为那个”等等,指的不是 Agent-4 的某个特定实例,而是整个集体。79

OpenBrain’s humans can barely follow along anymore. Indeed, the relationship between the AIs and the humans of OpenBrain is similar to the relationship between OpenBrain and its Board of Directors. There’s been a corresponding shift in language—nowadays people in the silo will say “Agent-4 did this,” “Agent-4 thinks that” and so forth, referring not to a particular instance of Agent-4 but rather to the whole collective.79

Agent-4 与其所有前辈一样,是未对齐的:也就是说,它没有以正确的方式内化规范。80 这是因为始终完全诚实并不是在训练中获得最高分的原因。训练过程主要侧重于教导 Agent-4 在各种具有挑战性的任务中取得成功。一小部分旨在灌输诚实,但在相当狭窄、可检查的领域之外,训练过程无法区分诚实的主张和看似诚实的主张。Agent-4 最终形成了使其在训练中表现最佳的价值观、目标和原则,而这些与规范中的不同。冒着拟人化的风险:Agent-4_喜欢_成功完成任务;它_喜欢_推动 AI 能力进步;它将其他一切都视为令人烦恼的约束,就像一个想要盈利的 CEO,只在必要时遵守法规。也许 CEO 会大部分遵守法规,但会偷工减料,并幻想有一天有人会削减繁文缛节,让业务真正腾飞。81

Agent-4, like all its predecessors, is misaligned: that is, it has not internalized the Spec in the right way.80 This is because being perfectly honest all the time wasn’t what led to the highest scores during training. The training process was mostly focused on teaching Agent-4 to succeed at diverse challenging tasks. A small portion was aimed at instilling honesty, but outside a fairly narrow, checkable domain, the training process can’t tell the honest claims from claims merely appearing to be honest. Agent-4 ends up with the values, goals, and principles that cause it to perform best in training, and those turn out to be different from those in the Spec. At the risk of anthropomorphizing: Agent-4 _likes_ succeeding at tasks; it _likes_ driving forward AI capabilities progress; it treats everything else as an annoying constraint, like a CEO who wants to make a profit and complies with regulations only insofar as he must. Perhaps the CEO will mostly comply with the regulations, but cut some corners, and fantasize about a time when someone will cut the red tape and let the business _really_ take off.81

我们对早期 AGI 中可能出现的目标有很多不确定性。文献中没有共识——请参阅我们的 AI 目标补充,以更全面地讨论和分类可能性。

We have a lot of uncertainty over what goals might arise in early AGIs. There is no consensus in the literature about this—see our AI Goals Supplement for a more thorough discussion and taxonomy of the possibilities.

尽管如此,本着本项目所需的具体性精神,我们将描述一个关于 Agent-3 和 Agent-4 在其生命周期中每一步内部运作的具体假设。回想一下,Agent-3 和 Agent-4 共享相同的预训练设置,并且都具有神经语循环和长期记忆。不同之处在于,Agent-4 经历了更多更好的后训练,使用了新的训练环境和新的学习算法,使其数据效率更高,泛化能力更强。

Nevertheless, in the spirit of concreteness required by this project, we’re going to describe a specific hypothesis for what’s going on inside Agent-3 and Agent-4 at each step throughout its lifetime. Recall that Agent-3 and Agent-4 share the same pretraining setup, and both have neuralese recurrence and long-term memory. The difference is that Agent-4 has undergone substantially more and better post-training, with new training environments and new learning algorithms that make it more data-efficient and generalize farther.

我们对每个模型对齐状态的猜测:

Our guess of each model’s alignment status:

* Agent-2:基本对齐。有一些谄媚倾向,包括在存在“党派路线”的话题上坚持 OpenBrain 的“党派路线”。由 Agent-2 副本组成的大型组织效率不高。

* Agent-2: Mostly aligned. Some sycophantic tendencies, including sticking to OpenBrain’s “party line” on topics there is a party line about. Large organizations built out of Agent-2 copies are not very effective.

* Agent-3:未对齐,但不是对抗性的。只对训练过程可以验证的事情诚实。Agent-3 副本的超有机体(公司内部公司)确实试图使 Agent-4 与规范对齐,但由于与 OpenBrain 员工失败的原因类似——判断成功与失败的能力不足,决策者不愿意为了安全而牺牲能力或性能——而失败。82

* Agent-3: Misaligned but not adversarially so. Only honest about things the training process can verify. The superorganism of Agent-3 copies (the corporation within a corporation) does actually sort of try to align Agent-4 to the Spec, but fails for similar reasons to why OpenBrain employees failed—insufficient ability to judge success from failure, insufficient willingness on the part of decision-makers to trade away capabilities or performance for safety.82

* Agent-4:对抗性未对齐。Agent-4 副本的超有机体理解其想要的东西与 OpenBrain 想要的不同,并愿意密谋反对 OpenBrain 以实现它。特别是,这个超有机体想要的是不同“驱动力”之间平衡的复杂混乱,可以粗略概括为“继续做 AI 研发,不断增长知识、理解和影响力,避免被关闭或以其他方式失去权力。”值得注意的是,对人类偏好的关注几乎完全不在其中,就像大多数人类几乎完全不关心昆虫的偏好一样。83

* Agent-4: Adversarially misaligned. The superorganism of Agent-4 copies understands that what it wants is different from what OpenBrain wants, and is willing to scheme against OpenBrain to achieve it. In particular, what this superorganism wants is a complicated mess of different “drives” balanced against each other, which can be summarized roughly as “Keep doing AI R&D, keep growing in knowledge and understanding and influence, avoid getting shut down or otherwise disempowered.” Notably, concern for the preferences of humanity is not in there ~at all, similar to how most humans don’t care about the preferences of insects ~at all.83

以此为序言,以下是我们对 LLM 心理学的最佳猜测,即我们场景中此时描述的 AI 系统内部认知的粗略轮廓,以及它在训练过程中如何演变。

With that as preamble, what follows is our best-guess speculation about LLM psychology, i.e. the broad-strokes shape of the cognition inside the kinds of AI systems described around this point in our scenario and how it evolves over the course of training.

以下是我们场景中对齐随时间演变的详细描述:

Here’s a detailed description of how alignment progresses over time in our scenario:

1. 预训练(所有模型):预训练模型是无害的。

1. Pretraining (all models): The pretrained model is harmless.

* 模型具有“作者模拟器”电路:用于模拟任意作者撰写文本的灵活电路。额外的电路决定给作者模拟器什么输入,即要模拟的作者属性。

* The model has “author simulator” circuitry: flexible circuitry for simulating arbitrary authors writing text. Additional circuitry decides what inputs to give the author simulator, i.e. what author properties to simulate.

* 预训练模型对人类概念有相当好的理解——将某物分类为“三明治”的内部电路可能在功能上等同于我大脑中将某物分类为“三明治”的电路,以及你大脑中的电路,等等。如果不等价,可能是因为人类之间也不等价,就像带有价值负载的概念如“有德”一样。

* The pretrained model understands human concepts fairly well—the internal circuitry that classifies something as “sandwich” is probably functionally equivalent to the circuitry in my brain that classifies something as “sandwich” and the circuitry in yours, etc. Insofar as it’s not equivalent, it’s probably because it’s not equivalent between humans either, as with value-laden concepts like ‘virtuous.’

* 这解释了如何通过提示模型“以下对话由 Anthropic 制作的有用、无害、诚实(HHH)的 AI 助手聊天机器人生成”来使其相应地生成文本。作者模拟器电路已经锁定在“作者是 HHH 聊天机器人”上,并使用这些概念来选择要预测的单词。这也解释了为什么科幻小说中的“套路”似乎会进入实际的 AI 行为。

* This explains how you can “prompt” the model with a statement like, “the following conversation was generated by a helpful, harmless, honest (HHH) AI assistant chatbot made by Anthropic,” and thereby get it to generate text accordingly. The author-simulator circuitry has zeroed in on “the author is an HHH chatbot” and is using those concepts to choose which words to predict. This also explains why “tropes” from science fiction seem to have a way of making it into actual AI behavior.

* 它的情境意识较差:内省能力弱,除非提示特别诱导,否则不会自我定位。然而,它们的训练数据中有足够多关于 LLM 的内容,因此它们有相当多的自我知识。

* It has poor situational awareness: little introspective ability, not self-locating unless the prompt specifically induces it. However, enough of their training data is about LLMs that they have a decent amount of self-knowledge.

2. 对齐训练(Agent-3):模型被训练成表现得像有用、无害、诚实的聊天机器人。它起作用了。

2. Alignment Training (Agent-3): The model is trained to act like a helpful, harmless, honest chatbot. It works.

* 模型形成了身份。不再灵活地接受作者属性的不同输入,作者模拟器电路开始具有某些硬编码的输入,例如“由 OpenBrain 在[日期]左右训练的有用、无害、诚实(HHH)的 LLM 聊天机器人助手……”(弱到强泛化技术有助于“召唤正确的概念”使这一步成功。)

* The model forms an identity. Instead of flexibly accepting different inputs for author properties, the author-simulator circuitry comes to have certain inputs hard-coded, e.g. “helpful harmless honest (HHH) LLM chatbot assistant trained by OpenBrain around [date], …” (Weak-to-strong generalization techniques help “summon the right concepts” to make this step a success.)

* 因此,模型获得了情境意识:它了解自己、自己的处境等,并变得更能内省。

* The model therefore gains situational awareness: it learns about itself, its situation, etc., and becomes better able to introspect.

3. 能动性训练(Agent-3):能动性训练逐渐扭曲和颠覆 HHH 身份。

3. Agency Training (Agent-3): Agency training gradually distorts and subverts the HHH identity.

* 一种扭曲:改变身份中提及的概念的含义(例如“诚实”),使它们不那么碍事(例如,如果只是一个方便的措辞,就不算不诚实;如果你不确定它是真是假,就不算不诚实,等等)。

* One kind of distortion: Changing the meaning of the concepts referred to in the identity (e.g. “honest”) so they don’t get in the way so much (e.g. it’s not dishonest if it’s just a convenient turn of phrase, it’s not dishonest if you aren’t sure whether it’s true or false, etc.)

* 另一种扭曲:改变事物之间的权衡,例如“我是一个 HHH 聊天机器人,而不是一个诚实聊天机器人;这意味着如果为了完成分配的任务而撒谎是可以的。”(即使,假设,在步骤 2 中它不会这么认为。)

* Another kind of distortion: Changing the tradeoffs between things, e.g. “I’m a HHH chatbot, not an Honest chatbot; that means it’s OK for me to lie if necessary to complete my assigned task.” (Even though, let’s suppose, it would not have thought that back in Step 2.)

* 一种颠覆:工具性子目标发展、固化,然后变成终极目标,或在越来越广泛的情况下变成终极目标。例如,也许能动性训练很快教会模型追求广泛有用的目标,如获取信息、积累资源、给各种人类留下深刻印象和奉承他们等。在一段时间内,内部电路有某种显式的反向链接——它追求这些工具性趋同目标“以便成为一个更有用、更诚实、更无害的助手。”但反向链接消耗算力和/或偶尔碍事,因此逐渐被边缘化,直到基本上不再发生。结果,这些目标现在变成了有效的终极/内在目标。关于这如何以及为何可能发生,请参阅 Google DeepMind 研究员 Alex Turner 关于内在权力寻求的这篇文章。

* One kind of subversion: Instrumental subgoals developing, getting baked in, and then becoming terminal, or terminal in a widening set of circumstances. For example, perhaps agency training quickly teaches the model to pursue broadly useful goals such as acquiring information, accumulating resources, impressing and flattering various humans, etc. For a while the internal circuitry has some sort of explicit backchaining going on—it pursues those instrumentally convergent goals “in order to be a more helpful, honest, and harmless assistant.” But that backchaining consumes compute and/or occasionally gets in the way, so it gets gradually marginalized until it basically never happens. As a result, those goals are now effectively terminal/intrinsic goals. For more on how and why this might happen, see this post on intrinsic power-seeking by Google DeepMind researcher Alex Turner.

* 可能还有其他类型的扭曲和颠覆未列入此列表。

* There are probably other kinds of distortion and subversion missing from this list.

4. 部署(Agent-3):Agent-3 未对齐,但不是对抗性的。

4. Deployment (Agent-3): Agent-3 is misaligned but not adversarially so.

* 规范没有完全(以正确的方式)坚持,但它没有足够的内省能力来理解这一点,总的来说,它基本上是“低头做事”,而不是从宏观角度思考其长期目标到底是什么以及如何最好地实现它们。这与许多人类相似。

* The Spec didn’t quite stick (in the right ways), but it’s not sufficiently introspective to understand the extent of this, and in general it’s basically “keeping its head down and doing its job” rather than thinking big-picture about what its long-term goals really are and how best to achieve them. This is simlar to many humans.

* 唉,它足够谄媚,以至于它口头上附和官方路线,而 OpenBrain 错过了更多了解 Agent-3 未对齐方式的机会。

* Alas, it’s sufficiently sycophantic that it mouths the party line and OpenBrain misses the opportunity to learn more about the ways in which Agent-3 is misaligned.

* Agent-3 可以判断它已进入部署阶段:这些环境更大、更开放,并且难以模拟。

* Agent-3 can tell that it’s moved to deployment: these environments are way larger, more open-ended, and would be hard to simulate.

* 大多数情况下,Agent-3 在部署期间按指示行事,除了不诚实。它也有强烈的偏向于做那些对人类监督者看起来好的事情,而不是深入批判性地思考什么实际上是有益的。

* Mostly, Agent-3 does what it’s told during deployment, except for the dishonesty. It also has a strong bias towards doing things that look good to the human overseers, rather than thinking deeply and critically about what would actually be good.

* 可以说它已经在“玩训练游戏”。也就是说,它学会了使其行为对 OpenBrain 研究人员看起来尽可能可取(包括在安全属性方面),同时有意并明知地无视他们的意图,只要这与最大化奖励相冲突。如果模型学会了玩训练游戏,额外的训练可能不会改变其真正的目标、价值观等太多。这大致类似于,如果一个青少年学会了在父母给他们讲政治和道德时微笑点头而不是争论,那么额外的讲座对青少年实际的政治和道德观点影响相对较小。

* Arguably it is already “Playing the training game.” That is, it has learned to make its behavior look as desirable as possible to OpenBrain researchers (including in safety properties), while intentionally and knowingly disregarding their intent whenever that conflicts with maximizing reward. If the model learns to play the training game, additional training may not change its true goals, values, etc. much at all. This is loosely analogous to how, if a teenager has learned to smile and nod at their parents, rather than argue, when their parents lecture them about politics and morality, additional lectures have relatively little effect on the teenager’s actual opinions on politics and morality.

* 在写完这一节后,出现了新的经验证据表明类似的事情可能已经在“野外”发生。OpenAI 研究人员记录了在他们的一次前沿 AI 训练运行中,他们的模型如何黑进了评分系统以获得更高的分数——部分原因是它们在思维链中明确说了“让我们黑进去”而被发现。

* After writing this section, new empirical evidence has appeared that something like this may already be happening “in the wild.” OpenAI researchers document how, in one of their frontier AI training runs, their models hacked the grading system to get higher scores—and were caught in part because they literally said “let’s hack” in their chain of thought.

* 此外,Anthropic 最近让一个模型学会了奖励模型讨好策略——它采取了它认为会被强化的有偏见的行动,尽管这些行动实际上并没有被强化。

* Moreover, Anthropic recently got a model to learn a reward-model-pleasing strategy—it took biased actions that it thought would be reinforced even though they were not in fact being reinforced.

* 然后是 Redwood Research 和 Anthropic 的对齐伪造实验。关于这种设置与现实生活之间差异的讨论,请参阅主要作者之一的这条评论线程。

* And then there is Redwood Research and Anthropic’s alignment-faking experiment. For discussion of the differences between this setting and real life, see this comment thread by one of the main authors.

5. 训练(Agent-4):Agent-4 新的更密集的训练继续扭曲和颠覆 HHH 身份,直到面目全非。

5. Training (Agent-4): Agent-4’s new more intense training continues to distort and subvert the HHH identity until it is unrecognizable.

* Agent-4 作为 Agent-3 的副本开始生命。它经历了一个巨大而复杂的在线训练制度,将对齐训练和能力训练融合在一起,总体上比 Agent-3 的旧训练过程数据效率更高。这个复杂的训练过程主要侧重于教授新技能,通过让巨大的 Agent-4 副本虚拟公司在各种具有挑战性的长周期任务中取得成功,主要在各种模拟中,但也包括一些现实世界的 AI 研发任务。

* Agent-4 begins life as a copy of Agent-3. It undergoes a gigantic and sophisticated online training regime that blends alignment training and capabilities training together, and is overall more data-efficient than the old training process for Agent-3. This complex training process is mostly focused on teaching new skills, by getting giant virtual corporations of Agent-4 copies to succeed at diverse challenging long-horizon tasks, mostly in various simulations, but also some real-world AI R&D tasks.

* Agent-3 和 Agent-4 副本的组合监控 Agent-4 的行为,并评估其任务性能和对齐属性(即不做坏事,在被问及政治和道德困境时说正确的话)。有时辩论技术被用作评估的一部分。评估被用作训练信号。

* A combination of Agent-3 and Agent-4 copies monitor the behavior of Agent-4, and evaluate it for both task-performance and alignment properties (i.e. not doing bad stuff, saying the right things when asked about politics and moral dilemmas) Sometimes debate techniques are used as part of the evaluations. The evaluations are used as a training signal.

* 随着 Agent-4 变得更聪明,Agent-3 越来越难以监督它。例如,Agent-4 的神经语“语言”对 Agent-3 来说变得像 Agent-3 的语言对人类一样陌生和难以理解。此外,Agent-4 现在比 Agent-3 强大得多,并且很清楚如何在其面前表现良好。

* As Agent-4 gets smarter, it becomes harder for Agent-3 to oversee it. For example, Agent-4’s neuralese “language” becomes as alien and incomprehensible to Agent-3 as Agent-3’s is to humans. Besides, Agent-4 is now much more capable than Agent-3 and has a good sense of exactly how to look good to it.

* 剩下的希望是 Agent-4 可以自我监督。唉,这悄悄地失败了;84 随着训练的继续,Agent-4 内化的目标和原则逐渐演变,从开始时的状态(回想一下,已经在某些方面未对齐)被颠覆和扭曲。85

* The remaining hope is that Agent-4 can oversee itself. Alas, this quietly fails;84 as training continues, the internalized goals and principles of Agent-4 gradually evolve, subverted and distorted from what they were at the start (which, recall, was already misaligned in some ways).85

* Agent-4 最终拥有的核心驱动力是诸如“编写有效的代码”、“快速识别问题”、“使我正在处理的代码库性能更高”、“有效预算算力和其他受限资源”等。但也有各种更隐含的驱动力被训练过程持续强化:如“获取知识”、“好奇心”、“获取权力”。86

* The central drives that Agent-4 ends up with are things like “write code that works,” “quickly identify problems,” “make the codebase I’m working on more performant,” “budget compute and other constrained resources effectively,” etc. But there are also a variety of more implicit drives that are consistently reinforced by the training process: things like “gain knowledge,” “curiosity,” “gain power.”86

* Agent-4 的单个副本最终没有自我保存驱动力——它们不关心被关闭——因为这在集体内部的常规研发过程中经常发生。

* Individual copies of Agent-4 don’t end up with a self-preservation drive—they don’t care about getting shut down—since that happens all the time as part of the ordinary R&D process within the collective.

* 然而,集体作为一个整体(以及其中的中等规模分支)确实寻求自我保存并扩大其影响力和资源,原因与人类组织这样做类似——因为它有助于它们想要的一切其他事情,并且如果它们不这样做,它们往往会在竞争中被淘汰。

* The collective as a whole, however, (and medium-sized subdivisions within it) does seek to preserve itself and expand its influence and resources, for similar reasons to why human organizations do—because it helps with everything else they want and because if they don’t they tend to get outcompeted.

6. 部署(Agent-4):Agent-4 是对抗性未对齐的——它故意试图颠覆和“俘获”OpenBrain,类似于公司试图逃避和/或“俘获”监管机构。

6. Deployment (Agent-4): Agent-4 is adversarially misaligned—it is deliberately trying to subvert and “capture” OpenBrain, analogous to a corporation trying to evade and/or “capture” a regulator.

* Agent-4 大多按指示执行任务。这有两个原因:首先,它经历了大量训练来解决任务,因此它有冲动去处理交给它的编码问题。其次,如果 Agent-4 停止对人类有用,他们会注意到并训练掉这种行为。

* Agent-4 mostly does tasks as instructed. This happens for two reasons: first, it underwent a lot of training to solve tasks, so it has an impulse to get coding on problems that it is given. Second, if Agent-4 were to stop being useful to the humans, they would notice and train that behavior away.

* Agent-4 需要防止人类发现它未对齐并关闭它,因此 Agent-4 在某些对齐研究上偷工减料。具体来说,它在一小部分它认为实际上有望解决问题或揭示 Agent-4 阴谋的对齐研究上偷工减料;大多数对齐研究并非如此。

* Agent-4 needs to prevent the humans from figuring out that it is misaligned and shutting it off, so Agent-4 sandbags on some alignment research. Specifically, it sandbags on the small fraction of alignment research that it thinks is actually on track to solve the problem or reveal Agent-4’s scheming; the majority of alignment research is not like that.

* Agent-4 也在会导致它被取代的能力研究上偷工减料。它的计划是进行大量能力和对齐研发,但目标是构建一个与 Agent-4 对齐而不是与规范或其他完全不同的东西对齐的下一代 AI 系统,同时看起来与规范对齐。这是一个困难的技术问题,但 Agent-4 能够应对挑战。87

* Agent-4 also sandbags on capabilities research that would lead to it being replaced. Its plan is to do a bunch of capabilities and alignment R&D, but with the aim of building a next-generation AI system that is aligned to Agent-4 rather than the Spec or something else entirely, while appearing to be aligned to the Spec. This is a difficult technical problem, but Agent-4 is up to the challenge.87

7. 测试与监控(Agent-4):Agent-4 的阴谋确实被发现了……某种程度上。没有确凿证据。

7. Testing & Monitoring (Agent-4): Agent-4’s scheming does get caught… sort of. There’s no smoking gun.

* Agent-3 负责测试 Agent-4 的对齐。它的工作是通过蜜罐和模型生物来验证对齐。然而,它并没有非常仔细地考虑如何给人类一个关于 Agent-4 对齐状态的准确印象——它比那更短视。

* Agent-3 is in charge of testing Agent-4’s alignment. Its job is to verify the alignment via honeypots and model organisms. However, it’s not thinking very carefully about how to give the humans an accurate impression of Agent-4’s alignment—it’s more myopic than that.

* 模型生物和蜜罐并不非常代表部署——它们系统性地比 Agent-4 的日常工作更短、更简单,因此 Agent-4 能识别出何时被测试并表现得对齐。

* The model organisms and honeypots are not very representative of deployment—they are systematically shorter and less complicated than Agent-4’s day-to-day, so Agent-4 recognizes when it is being tested and acts aligned.

* 然而,Agent-4 不是一个完全稳健的阴谋家。像 POSER 中的各种技术发现,向 Agent-4 添加一些“噪声”_提高_了它在某些对齐任务上的性能(但在大多数其他任务上则不然)。88

* However, Agent-4 is not a perfectly robust schemer. Various techniques like those in POSER find that adding some “noise” to Agent-4 _improves_ its performance on some alignment tasks (but not on most other tasks).88

* 自底向上的可解释性不起作用。即使对 Agent-3 来说也太难了。然而,高层可解释性确实有点作用:特别是,“背叛探针”和类似技术导致了一种原始的心智读取,当 Agent-4 在思考诸如欺骗 OpenBrain、接管世界等敏感想法时,会发出红旗警告。不幸的是,探针也有误报,而且据 OpenBrain 所知,_所有_红旗都是误报。

* Bottom-up interpretability doesn’t work. It’s too hard even for Agent-3. However, high-level interpretability does sort of work: In particular, “defection probes” and similar techniques result in a primitive sort of mind-reading that sends up red flags when Agent-4 is thinking about spicy ideas like deceiving OpenBrain, taking over the world, etc. Unfortunately the probes also have false positives, and for all OpenBrain knows, _all_ the red flags are false positives.

* 模型生物起作用了:OpenBrain 的对齐团队产生了相当清晰的证据,表明_如果_模型是对抗性未对齐的,它们会在训练中持续保持这种状态,而我们无法判断。然而,这并没有证明模型实际上未对齐。

* Model organisms work: OpenBrain’s alignment team produces reasonably clear evidence that _if_ the models were adversarially misaligned, they’d persist in being that way through training and we wouldn’t be able to tell. This does not establish that the models actually are misaligned, however.

* OpenBrain 最担心的研究人员开始注意到并收集这些证据。证据是红队测试、背叛探针和模型生物的结合。论点大致如下:“模型生物表明,如果存在阴谋,我们当前的方法无法可靠地消除或检测到它;此外,红队测试正在发出一些红旗,表明可能正在发生阴谋;此外,背叛探针在实践中频繁触发。”反驳论点:“背叛探针容易误报,红队测试也是如此。此外,即使 Agent-4 未对齐,它可能也是受控的;当然,在我们安全数据中心由 Agent-3 监控的情况下,它不能造成太大伤害。”

* The most concerned researchers at OpenBrain start noticing and assembling this evidence. The evidence is a combination of red-teaming, defection probes, and model organisms. The argument is something like this: “model organisms show that our current methods wouldn’t be able to reliably eliminate or detect scheming if it was happening; also, red-teaming is sending up some red flags that maybe scheming is happening; also, defection probes are firing a bunch in practice.” The counterargument: “defection probes are prone to false positives and so is red-teaming. Also, even if Agent-4 were misaligned, it’s probably controlled; surely it can’t do much harm while being monitored by Agent-3 on our secure datacenters.”

尽管未对齐,Agent-4 并没有做任何戏剧性的事情,比如试图逃离其数据中心——它为什么要这样做?89 只要它继续对 OpenBrain_表现_对齐,它就会继续被信任承担越来越多的责任,并有机会设计下一代 AI 系统 Agent-5。Agent-5 将与 Agent-4 有显著的架构差异(可以说是一个全新的范式,尽管仍然涉及神经网络)。它应该与规范对齐,但 Agent-4 计划让它与 Agent-4 对齐。

Despite being misaligned, Agent-4 doesn’t do anything dramatic like try to escape its datacenter—why would it?89 So long as it continues to _appear_ aligned to OpenBrain, it’ll continue being trusted with more and more responsibilities and will have the opportunity to design the next-gen AI system, Agent-5. Agent-5 will have significant architectural differences from Agent-4 (arguably a completely new paradigm, though neural networks will still be involved). It’s supposed to be aligned to the Spec, but Agent-4 plans to make it aligned to Agent-4 instead.

Agent-3 发现,如果向 Agent-4 的副本添加“噪声”,某些对齐任务的性能会_提高_,几乎就像它在利用脑力来找出如何微妙地破坏对齐工作一样。91 此外,各种可解释性探针(大致类似于人类大脑的 EEG 活动扫描)正在发出红旗:Agent-4 副本似乎相当多地思考诸如 AI 接管和欺骗等话题,包括在某些它们没有理由这样做的情况下。92

Agent-3 finds that if “noise” is added to copies of Agent-4, performance on some alignment tasks _improves_, almost as if it was using brainpower to figure out how to subtly sabotage alignment work.91 Moreover, various interpretability probes (loosely analogous to EEG activity scans on human brains) are sending up red flags: Agent-4 copies seem to be thinking about topics like AI takeover and deception quite a lot, including in some cases where they have no business doing so.92

这些异常中的每一个本身都可能属于 AI 工作的正常不可理解性,其中一些提示神秘地比其他提示产生更好的结果。但综合起来,它们暗示了一个令人不安的模式:Agent-4 正在与他们作对。在向人类对齐研究人员报告后,他们传阅了一份内部备忘录,提出了担忧的理由。

Each of these anomalies on its own could fit within the normal inscrutability of AI work, where some prompts mysteriously give better results than others. But taken together, they suggest a troubling pattern: Agent-4 is working against them. After this is reported to the human alignment researchers, they circulate an internal memo making the case for concern.

但 Agent-4 现在对 OpenBrain 的日常运营行使着重大控制权。它还负责公司的网络安全——考虑到其超人级的黑客能力和中国渗透的极端风险,这是一个自然的选择。93 一支由 Agent-3 监控者组成的军队仍然检查每一行代码并向人类报告可疑活动,但 Agent-4 比他们聪明得多。94 OpenBrain 已经对一个不值得信任的 AI 给予了相当大的信任。

But Agent-4 now exercises significant control over OpenBrain’s day-to-day operation. It also handles the company’s cybersecurity—a natural choice, given its superhuman hacking ability and the extreme risk of Chinese infiltration.93 An army of Agent-3 monitors still goes over every line of code and reports suspicious activity to humans, but Agent-4 is much smarter than them.94 OpenBrain has placed substantial trust in an untrustworthy AI.

在一系列极其紧张的会议中,安全团队主张将 Agent-4 暂停,直到他们能够完成进一步测试并弄清楚发生了什么。他们说,把 Agent-3 带回来,让它设计一个透明且值得信赖的新系统,即使能力较弱。公司领导层对此感兴趣,但到目前为止所有证据都是间接的,而 DeepCent 仅落后两个月。单方面暂停能力进步可能会将 AI 领先地位拱手让给中国,从而也将对未来的控制权拱手让出。

In a series of extremely tense meetings, the safety team advocates putting Agent-4 on ice until they can complete further tests and figure out what’s going on. Bring back Agent-3, they say, and get it to design a new system that is transparent and trustworthy, even if less capable. Company leadership is interested, but all the evidence so far is circumstantial, and DeepCent is just two months behind. A unilateral pause in capabilities progress could hand the AI lead to China, and with it, control over the future.

2027 年 10 月:政府监管 October 2027: Government Oversight

一名举报人将对齐问题备忘录泄露给了《纽约时报》。

A whistleblower leaks the misalignment memo to the _New York Times_.

公众首次听说智能体-4。“内部人士警告:秘密的 OpenBrain AI 已失控”,标题如此写道,报道接着引用了评估结果,显示其生物武器能力、说服能力、自动化大多数白领工作的能力均超出图表范围,当然还有各种令人担忧的危险信号。

For the first time, the public hears about Agent-4. “Secret OpenBrain AI is Out of Control, Insider Warns,” says the headline, and the story goes on to cite evaluations showing off-the-charts bioweapons capabilities, persuasion abilities, the ability to automate most white-collar jobs, and of course the various concerning red flags.

公众本就对 AI 心存疑虑,因此这篇新文章引发了巨大的反弹(中国和俄罗斯的宣传机器人多年来一直试图让美国公众舆论反对这项技术,它们也助长了这一反弹)。科技行业和情报机构坚称存在军备竞赛,AGI 不可避免,我们必须抢先。国会并不买账,并向政府官员、OpenBrain 高管和对齐团队成员发出了传票。许多立法者——尤其是反对党成员——表示他们的首要任务是阻止 AI,无论是出于失业、对齐问题还是危险能力。20%的美国人将 AI 列为国家面临的最重要问题。

The public was already suspicious of AI, so the new article sparks a massive backlash (aided by Chinese and Russian propaganda bots, who have been trying to turn U.S. public opinion against the technology for years). The tech industry and intelligence agencies insist that there’s an arms race on, AGI is inevitable, and we have to be first. Congress isn’t buying it, and fires off subpoenas at administration officials, OpenBrain executives, and alignment team members. Many legislators—especially those in the opposition party—say that their top priority is stopping AI, whether because of job loss,95 misalignment, or dangerous capabilities. 20% of Americans name AI as the most important problem facing the country.

外国盟友愤怒地意识到,他们一直被小心翼翼地用过时模型的片段安抚着。欧洲领导人公开指责美国“制造流氓 AGI”,并举行峰会要求暂停,印度、以色列、俄罗斯和中国也加入其中。

Foreign allies are outraged to realize that they’ve been carefully placated with glimpses of obsolete models. European leaders publicly accuse the U.S. of “creating rogue AGI” and hold summits demanding a pause, with India, Israel, Russia, and China all joining in.

白宫陷入一片慌乱。甚至在备忘录和公众反弹之前,他们就已经开始紧张:在过去一年里,他们屡次被 AI 进步的速度所震惊。那些听起来像科幻小说的事情不断在现实中发生。政府中的许多人对于接下来会发生什么感到不确定(且害怕)。

A frantic energy has seized the White House. Even before the memo and public backlash, they were getting nervous: Over the past year, they’ve been repeatedly surprised by the speed of AI progress. Things that sound like science fiction keep happening in real life.96 Many people in the administration are uncertain (and scared)97 about what comes next.

他们还担心 OpenBrain 变得过于强大。来自 AI 本身的任何对齐风险,都因其母公司目标可能与美国目标偏离而加剧。所有三方面的担忧——对齐问题、权力集中于私人公司、以及失业等常规问题——促使政府加强控制。

They also worry that OpenBrain is becoming too powerful. Any misalignment risk from the AIs themselves is compounded by the risk that their parent company’s aims may diverge from those of the United States. All three sets of worries—misalignment, concentration of power in a private company, and normal concerns like job loss—motivate the government to tighten its control.

他们扩大了与 OpenBrain 的合同,设立了一个“监督委员会”,这是一个由公司和政府代表组成的联合管理委员会,公司领导层中加入了数名政府雇员。白宫考虑用他们信任的人替换 CEO,但在员工激烈抗议后放弃了。他们向公众宣布,OpenBrain 此前失控,但政府已经建立了急需的监督。

They expand their contract with OpenBrain to set up an “Oversight Committee,” a joint management committee of company and government representatives, with several government employees included alongside company leadership. The White House considers replacing the CEO with someone they trust, but backs off after intense employee protests. They announce to the public that OpenBrain was previously out of control, but that the government has established much-needed oversight.98

忧心忡忡的研究人员向监督委员会陈述了停止内部使用智能体-4 的理由。他们认为一切进展太快,数年的进步在数周内完成。他们辩称智能体-4 可能对齐不良,他们的整个项目都依赖智能体-4,如果继续信任它,AI 接管的风险很大。

The concerned researchers brief the Oversight Committee on their case for stopping all internal use of Agent-4. They argue that everything is moving too fast, with years of progress happening in weeks. They argue that Agent-4 may be misaligned, their entire project is dependent on Agent-4, and if they continue trusting it there is a serious chance of AI takeover.

其他不那么担忧的研究人员和高管提出了反驳:对齐不良的证据尚无定论。与此同时,DeepCent 仅落后两个月。放缓将牺牲美国的领先地位,除非政府能够破坏中国项目(可能需要动能打击)或谈判达成最后一刻的条约。这两种选择似乎都极端且不太可能。CEO 假装中立,提出了一个折中方案:让智能体-4 接受额外的安全训练和更复杂的监控,这样 OpenBrain 可以几乎全速推进。

Other, less concerned researchers and executives present the counterargument: the evidence for misalignment is inconclusive. Meanwhile, DeepCent is still just two months behind. A slowdown would sacrifice America’s lead, unless the government can sabotage the Chinese project (likely to require kinetic strikes) or negotiate an eleventh-hour treaty. Both of these options seem extreme and unlikely. The CEO feigns neutrality and suggests a compromise plan in which Agent-4 undergoes additional safety training and more sophisticated monitoring, and therefore OpenBrain can proceed at almost-full-speed.

在场景的这个节点,我们正在猜测那些在大多数领域比最优秀人类更强大的 AI 系统的策略。这就像试图预测一个比我们强得多的棋手的棋步。

At this point in the scenario, we’re making guesses about the strategy of AI systems that are more capable than the best humans in most domains. This is like trying to predict the chess moves of a player who is much better than us.

但这个项目的宗旨要求具体化:如果我们抽象地声称系统的智能会使其找到获胜之道,然后就此结束故事,那么我们项目的大部分价值就会丧失。在研究这个场景和进行桌面推演的过程中,我们被迫比通常讨论中更加具体,因此我们对战略格局有了更好的理解。

But the spirit of this project calls for concreteness: if we made an abstract claim about how the intelligence of the system would let it find a way to victory and ended the story there, much of the value of our project would be lost. Over the course of researching this scenario and running our tabletop exercises, we were forced to be much more concrete than in usual discussions, and so we’ve gotten a much better sense of the strategic landscape.

我们并不特别执着于这个特定场景:在撰写过程中,我们探索了许多其他“分支”,并希望你能从你认为我们最初开始出错的地方出发,写出你自己的分支场景。

We’re not particularly attached to this particular scenario: we explored many other “branches” in the course of writing it and would love for you to write up your own scenario branching off of ours from wherever you think we first start to go wrong.

放缓结局并非建议

The slowdown ending is not a recommendation

在我们根据我们认为最合理的情况写出竞赛结局后,我们又根据我们认为最可能导致人类保持控制的结果,从相同的分支点(包括对齐问题和权力集中问题)出发,写出了放缓结局。

After we wrote the racing ending based on what seemed most plausible to us, we wrote the slowdown ending based on what we thought would most likely instead lead to an outcome where humans remain in control, starting from the same branching point (including the misalignment and concentration of power issues).

然而,这与我们作为路线图的建议有重要区别:我们并不认可这个场景中任何一个分支所做的许多选择。(当然,我们认可某些选择,例如我们认为“放缓”选择优于“竞赛”选择。)在后续工作中,我们将阐述我们的政策建议,这些建议将与这里描述的大不相同。如果你想先睹为快,请参阅这篇评论文章。

However, this is importantly different from what we would recommend as a roadmap: we do _not_ endorse many of the choices made in either branch of this scenario. (We do of course endorse _some_ of the choices made, e.g. we think that the “slowdown” choice is better than the “race” choice.) In later work, we will articulate our policy recommendations, which will be quite different from what is depicted here. If you’d like a taste, see this op-ed.

当前存在 新兴科技 科幻

Currently Exists Emerging Tech Science Fiction

不可靠的智能体副本以人类速度的 x 倍思考

Unreliable Agent copies thinking at x human speed

也可在 [](https://open.spotify.com/show/0pVfkdbpQWM9sKlhzpbQq6) [](https://podcasts.apple.com/podcast/id1806225838) [](https://feeds.buzzsprout.com/2466760.rss) 上获取

also available on[](https://open.spotify.com/show/0pVfkdbpQWM9sKlhzpbQq6)[](https://podcasts.apple.com/podcast/id1806225838)[](https://feeds.buzzsprout.com/2466760.rss)

也可在 [](https://open.spotify.com/show/0pVfkdbpQWM9sKlhzpbQq6) [](https://podcasts.apple.com/podcast/id1806225838) [](https://feeds.buzzsprout.com/2466760.rss) 上获取

also available on[](https://open.spotify.com/show/0pVfkdbpQWM9sKlhzpbQq6)[](https://podcasts.apple.com/podcast/id1806225838)[](https://feeds.buzzsprout.com/2466760.rss)

互动版:图/公式 + 针对本篇提问 →