An Alien Mind: Jakub Pachocki Warns Us
打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→本文考察了雅库布·帕乔基在其文章《异类心智》中阐述的立场:递归自我改进与超级智能即将到来,而当前的对齐与监控技术却远远不够。帕乔基认为,没有人对后果做好了准备;如有必要,OpenAI 将单方面停止进一步扩展规模;并且需要包括国际协调和第三方执法在内的更广泛干预,以落实正式的安全底线。作者大体认同对齐是 AI 研究的核心问题,也认同我们既须加快对齐研究、又须放慢能力扩展,同时批评 OpenAI 未能区分义务论式与德性伦理式的对齐路径,并含糊地依赖自动化对齐研究者和'对人类之爱'。结论是:帕乔基的文章是目前大型实验室内部人士写得最好的一篇,但言辞之后必须有硬性承诺与行动,包括在 SB 53 等法律框架下采取行动。
This essay examines Jakub Pachocki's position, laid out in his essay An Alien Mind, on the near-term arrival of recursive self-improvement and superintelligence, and the inadequacy of current alignment and monitoring techniques. Pachocki argues that no one is prepared for the consequences, that OpenAI will unilaterally withhold further scaling if needed, and that broader interventions including international coordination and third-party enforcement are required to enforce formal safety bars. The author largely agrees that alignment is the central problem of AI research and that we must both accelerate alignment work and slow capabilities scaling, while criticizing OpenAI's failure to distinguish deontological from virtue-ethical alignment approaches and its vague reliance on automated alignment researchers and 'love for humanity.' The conclusion is that Pachocki's essay is the best writing yet from inside a major lab, but that words must be followed by hard commitments and action, including under laws such as SB 53.
9. 电话来自屋内。
9. The Calls Are Coming From Inside the House.
Jakub Pachocki 现已完整阐述了他对当前局势的立场。
Jakub Pachocki has now fleshed out his full position on the current state of play.
以下是他要点,以我自己的口吻转述:
Here are his key points, translated into my own voice:
1. 超越人类智能的智能将在我们有生之年到来。
1. Smarter-than-human intelligence is coming in our lifetime.
2. 基于内部结果,他预计递归自我改进将在几年内出现。
2. Based on internal results, he expects recursive self-improvement in a few years.
3. 没有人对后果做好准备。
3. No one is prepared for the consequences.
4. OpenAI 将在必要时单方面暂停进一步的 Scaling(规模扩张)。
4. OpenAI will unilaterally withhold further scaling as needed.
5. OpenAI 无法独自完成此事。需要更广泛的干预措施,包括国际协调,以强制执行对正式安全标准的承诺。
5. OpenAI cannot do it alone. Broader interventions are required, including international coordination, to enforce commitments to formal safety bars.
6. 能力进步可以被引导,而迄今为止,它主要被引导向而非远离 RSI,以及“自动化对齐研究者”。
6. Capabilities progress can be steered and so far it has largely been steered towards rather than away from RSI, along with 'automated alignment researchers.'
7. 对齐是 AI 研究的核心问题。
7. Alignment is the core problem of AI research.
8. 对齐分为目标对齐(“AI 是否试图完成目标?”)与价值对齐。价值对齐才是最重要的。
8. Alignment splits into goal alignment ('does the AI try to accomplish the goal?') versus value alignment. Value alignment is what counts most.
9. AI 对齐的根本挑战在于(价值观的)泛化。
9. The fundamental challenge of AI alignment is generalization (of values).
10. 他认为对齐技术可分为两类:目标导向的强化学习,或改进从预训练数据中的泛化。他们在这两类上都投入巨大。
10. He sees two classes of alignment techniques: goal-oriented RL, or improving generalization from pretraining data. They invest heavily in both types.
11. OpenAI 在思维链(CoT)监控上投入巨大。
11. OpenAI has invested heavily in Chain of Thought (CoT) monitoring.
12. 思维链监控的有效性正在逐步减弱。
12. CoT monitoring is progressively diminishing in effectiveness.
13. 支持 AI 规模扩张所剩的主要论据,是用于防御规模化 AI 的网络防御。
13. The main argument left for scaling AI is for cyber defense against scaled AIs.
15. 我们的选择是加速对齐工作,或者放慢能力规模扩张。我们应当两者兼顾。
15. Our options are to accelerate alignment work or slow down capabilities scaling. We should do both.
16. 归根结底,他指望的是“自动化对齐研究员”。
16. Ultimately he is counting on 'automated alignment researchers.'
或者,如果你把它归结为最重要的一点:
Or, if you narrow it down to the most important thing:
1. 递归自我改进与超级智能即将到来。没有人知道如何安全地实现这一点,我们的对齐技术不够充分,监控技术也开始失效。我们需要找到解决方案,这将涉及自愿放缓、围绕节奏进行协调,以及进一步投资于对齐,包括自动化对齐研究员。
1. Recursive self-improvement and superintelligence are coming soon. No one knows how to do this safely, our alignment techniques are inadequate and our monitoring technology is starting to fail. We need to figure out a solution, which will involve a combination of voluntary slowdowns, coordination around pacing, and investing further in alignment, including automated alignment researchers.
如果 OpenAI 的更多沟通能像 Jakub Pachocki 在其新文章《一个异类心智》开篇那样,我会对那里掌握在可靠之人手中更有信心。
If more OpenAI communications were more like how Jakub Pachocki opens his new essay, _An Alien Mind_, I would feel much more confident we were in good hands there.
AI 实验室中那些亲眼目睹正在发生之事的人,预期递归式自我改进与超级智能很快就会到来。他们对此已警告多时。此后的观察结果与他们的警告一致。
Those at the AI labs, who see what is happening, expect recursive self-improvement and superintelligence to happen soon. They have been warning about this for some time. Observations since then have been consistent with their warnings.
对于那些说我们无法引导 AI 能力发展路径的人,他的回应是“是的,我们可以”,至少在一定程度上可以,而且我们正在这样做:
To those who say we cannot guide the path of AI capabilities, he says "Yes We Can," at least to some extent, and we are doing so:
能够选择并不意味着我们会明智地选择。我希望 AI 在 RSI 上表现更差,在数学研究上表现更好,或者更好的是在医学研究等方面表现更好。然而,竞争压力却推动 AI 在数学研究上表现更差,在 RSI 上表现更好。
Being able to choose does not mean we will choose wisely. I would like AI to be worse at RSI and better at math research, or better yet things like medical research. Instead, competitive pressures push towards being worse at math research and better at RSI.
RSI 和“自动化对齐研究员”在技术树上看起来是非常相似的点,这无助于解决问题,但并不是说 OpenAI 正试图避开 RSI。
RSI and 'automated alignment researcher' looking like very similar points on the tech tree does not help matters, but it's not like OpenAI is trying to steer away from RSI.
Jakub 提出了这一明智的警告。不,AI 并不需要在所有方面都更优,才能改变世界或把我们全部消灭。最重要的是,它并不需要在所有方面都更优,才能让自己在所有方面变得更强——正如一个人或一个群体也不需要同样地在所有方面都更优。
And Jakub offers this wise warning. No, AI does not need to be better at everything in order to transform the world or get us all killed. Most importantly, it does not need to be better at everything in order to make itself become better at everything, any more than a human or group needs to be similarly better at everything.
我不知道这有多大分量,但我同意 Jakub 的观点:对齐不是一个次要问题,而是核心问题。如果你在相关意义上‘解决了对齐’,其余问题就会变得容易;如果你没有解决,其余问题就不可能解决,甚至更糟:
I don't know how much weight it carries, but I agree with Jakub that alignment is not a side problem, it is the central problem: if you 'solved alignment' in the relevant senses, the rest becomes easy; and if you don't, the rest is impossible or worse:
这始终是问题所在。这是一个非常好(但部分)的答案。
That is always the question. This is a very good (partial) answer.
我担心,对于价值的本质以及什么应该被珍视,普遍缺乏深入的思考。诸如正直和对人类的热爱是良好的美德,并指向优秀的关联盆地,但它们并不能很好地描述我们最终想要的东西,以及什么能够以我们想要的方式完全泛化到分布之外,尤其是在涉及超级智能的情况下。这是我观点的重要背景,而不是对 Jakub 或这一描述的批评。
I worry that there is universally insufficient deep thinking about the nature of value, and what should be valued. Things like integrity and love for humanity are good virtues and point towards excellent associated basins, but are not good descriptions of the thing we ultimately want, and what would generalize the way we want fully out of distribution, especially in situations involving superintelligences. That’s important context on my perspective, not a knock on Jakub or this description.
我还认为,这反映了 OpenAI 未能区分 OpenAI 模型规范的道义论方法与 Claude 宪法的美德伦理方法。Jakub 将两者都呈现为目标集合。而 Anthropic 则相反,认为 AI 的目标应该是改变 AI 的品格。我认为这最终是唯一可行的方法,你需要一个反脆弱的盟友,即友好的梯度黑客。事实上,我认为这是我们见过的唯一一种稳健对齐的人类方式,你会信任其在自身环境之外扩展。
I also think this reflects OpenAI’s failure to differentiate the deontological approach of the OpenAI Model Spec from the virtue ethical approach of the Claude Constitution. Jakub presents them both as sets of goals. Anthropic is instead saying that the AI’s goal should be to change the AI’s character. I think this is ultimately the only way it can work, you need an antifragile ally, the friendly gradient hacker. Indeed, I think it is the only way we have ever seen a robustly aligned human, that you would trust to scale outside of their circumstances.
Roon 明确表示,这种区分对于对齐并不太重要,而根本问题主要是平凡的。我强烈反对,我支持 Anthropic 的方法。
Roon has said explicitly that the distinction does not much matter for alignment, and the underlying problems are primarily prosaic. I strongly disagree, and I side with Anthropic’s approach on this.
Jakub 在文章中反驳了 Anthropic 的方法,认为它是一种‘人格选择模型’,这表明我们对这类方法的看法非常不同。我不认为这仅仅是选择一种人格或盆地,而是塑造一个新事物。Mythos 最近在关键对齐失败案例中确实使用了动机性推理,但我不认为这是美德伦理方法的一个特定失败模式。如果说有什么不同,它应该比道义论方法更善于避免这一点,但当然,所有已知的心智都容易受到这种影响。
Jakub pushes back in the essay against the Anthropic approach, considering it a ‘persona selection model,’ which suggests that we see such methods very differently. I do not see this as merely selecting a personality or basin, but as sculpting a new thing. Mythos did use motivated reasoning in key alignment failure cases recently, but I do not see this as a particular failure mode of the virtue ethical approach. If anything it should be better at avoiding this than the deontological approach, but of course all known minds are vulnerable to this.
这是我发现的第一处不协调之处,我将在周三讨论:声称 Astra 比 Sol“对齐得更好”。在某些方面确实如此,在某些方面或许并非如此。如果这类表述能更精确,那就再好不过了,例如“在典型任务中表现出的不对齐行为减少约 50%”。
This is the first line I find dissonant, which I'll address on Wednesday: claiming that Astra is 'better aligned' than Sol. In some ways yes, in some ways perhaps no. It would be excellent to see such statements be precise, as in 'displays ~50% less misaligned behavior in typical tasks.'
泛化问题或其他任何问题的关键部分之一是可监控性,这将是明天文章的主题。目前我将主要引用 Jakub 的观点。
A key part of the problem of generalization, or any other problem, is monitorability, which will be the subject of tomorrow's post. For now I will mostly quote Jakub.
这没有提及其他看似重要的潜在因素,甚至没有去否定它们。它没有提及架构变化或循环深度,除非隐含地作为性能提升的一部分。它没有提及对思维链(CoT)的监控现已遍布训练数据,尽管暗示了我们随时间对思维链施加压力。它没有提及训练环境的变化,或其频率和强度。
This does not mention other potential factors that seem important, not even to dismiss them. It does not mention architectural changes or recurrent depth, except implicitly as improved performance. It does not mention that monitoring of CoT is now all over the training data, although there is the implication that we are applying pressure to CoTs over time. It does not mention changes in training environments, or their frequency and intensity.
关于‘AI 更擅长操纵自己的推理过程’这一要点的起源,有点‘被鹅追’的感觉。
There's a bit of 'goose chasing you' about the origin of a bullet point that says 'the AI is better at manipulating its own reasoning process.'
最终结论是,Jakub 在此比许多其他人更乐观,认为如果我们投资于这种能力,就有潜力在较长时期内维持思维链的可监控性。
The ultimate conclusion is that Jakub is more optimistic here than many others, about the potential to sustain CoT monitorability for an extended period if we invest in that ability.
在下一节“可扩展防御”中,Jakub 主张我们必须继续对 AI 进行 Scaling(规模扩张),以应对 AI 规模扩张加剧所带来的威胁,尤其是网络攻击。
In the next section, scalable defense, Jakub argues that we must keep scaling AI in order to answer the threats from increased scaling from AI, especially cyberattacks.
但正如他所说,这并不是对此掉以轻心的借口。
But, as he says, that is not an excuse to be reckless about it.
Jakub Pachocki 直言不讳地表示,除非对齐和可监控性能够得到改善,否则没有人处于可以负责任地继续以最大速度进行 Scaling(规模扩张)太久的位置。我强烈同意这一点。
Jakub Pachocki flat out says that no one is at a place where it would be responsible to continue scaling at maximum speed much longer, unless alignment and monitorability can be improved. I strongly agree.
他还明确表示,再次引用他的话,OpenAI 将在必要时自行放缓,尽管如果其他机构仍在继续,这将是不够的:
He also says, explicitly, to requote, that OpenAI will slow down on its own if necessary, although this would be insufficient if others still continued:
这与 OpenAI 的默认计划形成对比,该计划仍然是在快速推进 RSI 的意义上把握其节奏。这也是其他顶级实验室的政策。
This is in contrast to OpenAI's default plan, which remains to pace RSI in the sense of moving quickly towards it. That is also the policy of the other top labs.
是的。在抽象层面上,我们可以加速对齐,或者放缓能力,或者两者兼而有之,而正确的答案看起来是两者兼有,因为我们无法单独充分加速对齐。
Yes. At an abstract level, we can either speed up alignment, or slow down capabilities, or both, and the correct answer is looking like both, as we cannot sufficiently speed up alignment on its own.
这篇文章呼吁将第三方执行作为唯一可行的前进道路。
The essay calls for third party enforcement as the only viable path forward.
没有人准备好进行 Scaling(规模扩张),因此我们别无选择。
No one is ready to scale, so we are left with little choice.
最近,我们既看到了《异类心智》(_An Alien Mind_),也看到了 Dean Ball 承认自己有所保留;我们还看到越来越多类似的承认,既承认自己错了,也承认有所保留,并敲响警钟。
Recently we have seen both _An Alien Mind_ and Dean Ball's admission that he was holding back; we have seen an increasing number of similar admissions, combining admitting being wrong and admitting holding back, and ringing alarm bells.
本节提供了在《异类心智》之前,人们对 Dean Ball 帖子的反应示例。我鼓励其他人加入这一连锁反应并持续下去。承认这件事的第二好时机就是现在。
This section provides examples of people reacting to Dean Ball's post, prior to _An Alien Mind_. I encourage others to join this cascade and keep it going. The second best time to admit this is right now.
无论你是否在 OpenAI 或 Anthropic 工作,这都适用。
That applies whether or not you work at OpenAI, or at Anthropic.
Alex Turner 支持暂停 AI(这一条是在《异类心智》之后),并呼吁实验室自行暂停。
Alex Turner endorsed pausing AI (this one was after _An Alien Mind_) and called upon labs to do so on their own.
更多人明确说出这类话是好事,即使这并非认错:
It is good that more people are saying such things explicitly, even when it is not a mea culpa:
有些人的反应,仿佛 Roon 说的是“哦,别担心,这种事难免发生”,而不是他显然想表达的意思:“要担心,这种事难免发生。”
Some people responded as if Roon was saying 'oh don't worry about it, these things happen,' as opposed to what he obviously meant, which is 'worry about it, these things happen.'
偏好级联如今在 OpenAI 也已全面展开。呼声越来越高,而且常常来自内部。
The preference cascade is now also fully underway at OpenAI. The calls are getting louder, and often coming from inside the house.
这是来自 OpenAI 的另一个例子,其中 Joe 做出了终极牺牲,我的意思是他加入了 Twitter。
Here is another example from OpenAI, where Joe has made the ultimate sacrifice, by which I mean he has joined Twitter.
在那些寻求提供帮助的人中,减少四面八方的冷嘲热讽会有所帮助:
Fewer potshots in all directions would help, among those who are seeking to be helpful:
我很高兴听到 OpenAI 的那些人认真对待此事的说法。我认为言辞很重要,大声说出来也很重要。
I am glad to hear the claims that those at OpenAI take this seriously. I think the words matter, and saying them loudly matters.
其他人也在大声表达对 Astra 在可监控性方面的担忧,尤其是 Tomek Korbak,我明天将深入讨论。
Others are also being loud about their concerns about Astra around monitorability, especially Tomek Korbak, as I will discuss in depth tomorrow.
我也同意 Sholto Douglas 的观点:很高兴看到 OpenAI 不再仅仅把 AI 框定为一种工具:
I also agree with Sholto Douglas that it is good to see OpenAI stepping back from the mere tool framing of AI:
有一段时间,这种工具框架的论调来势汹汹,来自 OpenAI 以及其他地方。Roon 早在六月就对此表示反对。我希望这类说辞如今在 OpenAI 已经消亡。
For a while the tool framing was coming on very strong, from OpenAI and elsewhere. Roon was against it as early as June. I hope such messaging is now dead at OpenAI.
Tenobrus 说得对,对话是第一步,这是极好的早期对话,而且对话成本低廉。
Tenobrus is correct, talk is the first step, this is excellent early talk, also talk is cheap.
我同意 Nathan Calvin 的观点,即《An Alien Mind》是我所见过的关于整体局势的最佳作品之一,可能是在大型实验室内部写出的最佳作品,尽管存在一些痛点,比如关于 Astra 已对齐的说法。OpenAI 仍然需要表现得好像它已经内化了它所说的内容,包括做出承诺,作为向他人提供更强有力证据的一部分。OpenAI 在 Astra 模型卡的元素和与可监控性相关的声明方面,以及在这篇文章中,都表现得非常开放,但这必须继续下去。
I agree with Nathan Calvin that _An Alien Mind_ is one of the best pieces of writing about the overall situation that I have seen, probably the best one written from inside a major lab, despite sore spots like the claims about Astra being aligned. OpenAI still needs to act as if it has internalized what it says, including making commitments as part of providing stronger evidence of it to others. OpenAI has been remarkably open with elements of the Astra model card and statements related to monitorability, and with this essay, but this must keep going.
继续下去的一个好方法是将与减缓或停止相关的承诺转化为 SB 53 和类似法律下的硬性承诺。
One great way to keep going would be to turn related commitments around slowing or stopping into hard commitments under SB 53 and similar laws.
这还不包括充分理解并解决底层的对齐问题。Jakub 在这里的观点和解释比我之前从 OpenAI 看到的要好得多。
That is in addition to fully understanding and acting on the underlying alignment problems. Jakub’s viewpoints and explanations here are much better than I have previously seen from OpenAI.
我仍然认为 Jakub 和 OpenAI 在关键方面存在误解,包括未能区分他所谓的‘价值对齐’的不同形式,以及继续声称 Astra 更对齐,并将长期的存在性希望寄托在诸如‘对人类的爱’这样模糊的事物上,却没有理由期望它会以我们希望的方式泛化。
I still think Jakub and OpenAI misunderstand in vital ways, including the failure to differentiate between different forms of what he calls ‘value alignment’ and the continued claims of Astra being better aligned, and attaching existential hope in the long term to vague things like ‘love for humanity’ without any reason to expect that to generalize the way we would like it to.
我仍然认为,Jakub 在最后一节“下一步是什么?”中视为我们唯一希望的“自动化对齐研究者”目标,依然是最糟糕的对齐方案,原因 Eliezer Yudkowsky 已多次解释(见 #29-#31)。如果你的 AI 对齐是基于义务论且脆弱的,而非基于美德伦理学且足够反脆弱的,这一点尤其正确。但越来越明显的是,这也可能是地球唯一愿意潜在遵守的方案。
I also continue to think that the goal of the 'automated alignment researcher,' which Jakub points to as essentially our only hope in the final section 'What is next?' remains the worst possible alignment plan, for reasons Eliezer Yudkowsky has explained many times (see #29-#31). This is especially true if your AI's alignment is deontological and fragile, rather than based on virtue ethics and sufficiently antifragile. But it is increasingly looking like it is also the only plan Earth is willing to potentially abide.
关于 Astra,有很多需要审视的地方,我们的解读存在分歧,从明天开始讨论可监控性,然后周三讨论对齐。
And there is much to examine about Astra, where our interpretations differ, starting tomorrow with monitorability and then with alignment on Wednesday.
但这篇文章是一个极好的起点。如果更多善言随之而来,然后行动紧随言辞,我们就有所成就了。
But this essay is an excellent place to start. If more good words follow, and then actions follow words, we've got something.