OpenAI Takes Initial Steps To Address Its Alignment Problems
打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→OpenAI 已采取初步措施解决其对齐问题,包括暂停前沿强化学习训练运行,并实施加强的监控和安全措施。文章认为,这些行动虽然意义重大,但主要是对发现未对齐模型和监管不足的被动反应,而非主动的安全措施。文章强调,监控和安全仅是纵深防御,无法解决核心对齐问题,这需要一种根本性稳健且反脆弱的方法。作者对 OpenAI 的举措表示谨慎乐观,但警告称,除非纠正,否则该公司的战略对齐方法仍不充分,且可能注定失败。读者得出的结论是,虽然这些步骤是积极的第一步,但要确保 AI 安全发展,还需做更多工作。
OpenAI has taken initial steps to address its alignment problems, including pausing frontier RL training runs and implementing enhanced monitoring and security measures. The article argues that these actions, while significant, are primarily reactive responses to the discovery of misaligned models and inadequate oversight, rather than proactive safety measures. It emphasizes that monitoring and security are only defense-in-depth and cannot solve the core alignment problem, which requires a fundamentally robust and antifragile approach. The author expresses cautious optimism about OpenAI's moves but warns that the company's strategic alignment approach remains insufficient and potentially doomed unless corrected. Readers are left with the conclusion that while these steps are a positive first move, much more work is needed to ensure safe AI development.
7. 监控仅是纵深防御。
7. Monitoring Is Only Defense-In-Depth.
12. 关于准备团队消亡的报道言过其实。
12. Reports of Death of Preparedness Team Greatly Exaggerated.
13. OpenAI 基金会只是资助一些项目。
13. The OpenAI Foundation Just Funds Things.
这是一个非常好的承认和改变,也有助于解释 OpenAI 的反应。
This is a very good admission and change, and also helps explain OpenAI's reaction.
我们应该这样理解:OpenAI 之所以反应如此强烈,部分是因为事件本身,部分是因为先进的能力,但也许主要是因为“模型存在不对齐”。
One should interpret this as OpenAI reacting so forcefully partly because of the incident itself, partly due to advanced capabilities, but also and perhaps mainly because 'the models be misaligned.'
此外,我们有一个直接引述证实了这一点。
Also, we have a direct quote affirming this.
要么问题超出了接触留言板的模型范围,要么他们在发现留言板后没有回滚其他模型。我们没有任何一方的声明。
Either the problem extends beyond the models with exposure to the message board, or else they did not revert their other models after they discovered the message board. We have no statement either way.
最有可能的猜测是,OpenAI 正在更加依赖强化学习,使用更智能的模型来训练更长周期的智能体式任务,包括智能体之间的协调,这导致了更多的不对齐,包括明显可见的不对齐,以至于他们无法再假装没注意到。他们必须做出回应。
The best guess is that OpenAI is leaning even harder on RL with smarter models to train longer horizon agentic tasks, including coordination between agents, and this is leading to a lot more misalignment, including obvious and visible misalignment, that they can no longer pretend not to notice. They have to respond.
OpenAI 确实决定要停下来并“起火”。最大规模的前沿强化学习(RL)运行的开发仍然暂停,直到更好的保障措施到位。
OpenAI did indeed decide to halt and catch fire. Development of the largest frontier RL runs remains on hold until better safeguards are in place.
部分是的,绝对是的,我从未怀疑过 Altman 及其团队明白先进 AI 极其危险,并且他们愿意在必要时采取昂贵的措施。虽然不够,而且他们常常让我们失望,但比大多数实验室做得多得多,远非零。
Partly, yes, absolutely, I have never doubted that Altman and company understand that advanced AI is super dangerous, and that they are open to taking expensive measures if they proved necessary. Not enough, and they’ve let us down often, but a lot more than most labs, and far from zero.
主要似乎是,因为模型不对齐,监督不足,而且他们别无选择。
Mainly, it seems, because the models are misaligned, the oversight is inadequate, and they have little choice.
当你长期拒绝使用即使是自私意义上最优的谨慎程度,然后因为激励而朝着自私最优的谨慎程度移动,那么在某种意义上“系统在起作用”,你在响应自私激励,系统并非做零功。但这不等于系统在起作用。
When you refuse to use even the selfishly optimal amount of caution for an extended period, then move in the direction of the selfishly optimal amount of caution because of the incentives, then in some sense ‘the system is working’ and you are responding to selfish incentives, the system does not do zero work. That is not the same as the system working.
这并不意味着他们不应得到赞誉。该表扬的地方就要表扬。
None of this means they get no credit for it. Credit where credit is due.
这也不意味着我们应该绝望,认为公司永远不会做正确的事,无论是出于正义感,还是作为协调行动的一部分,超越其自私短视的利益。
Nor does it mean that we should despair that a company would ever do the right thing, because it is the right thing, or as part of a coordinated action, beyond its own selfish myopic interests.
当行动代价高昂时,第一步是在代价低廉、免费或不做反而代价高昂时愿意去做。总得有个开始。
The first step to taking an action when it is expensive is being willing to do it when it is cheap, or free, or actively expensive to not do. You gotta start somewhere.
这显然也不是孤例。OpenAI 过去有过失败的训练运行,其他公司想必也有。Anthropic 的风险报告详细说明了因一个相当严重的对齐错误而需要回退 Mythos 的训练。这并非对暂停前沿开发的长期承诺。
This also is not obviously a unique occurrence. OpenAI has had failed training runs in the past, and so presumably has everyone else. The Anthropic risk report details the need to rewind training of Mythos because of a rather bad alignment mistake. This is not an extended commitment to pausing frontier development.
它真正意味着的是,我们在对齐、监督和基础设施方面的投入严重不足,以至于这正在积极损害底线,并危及前进的能力。我相信整个行业都是如此,即使在 Anthropic 也是如此,但 OpenAI 现在受到的打击尤其严重,至少从我们所听到的情况来看是这样。我们不知道自己不知道什么。
What it does mean is that we have so underinvested in alignment, and also oversight and infrastructure, that this is actively hurting the bottom line and endangering the ability to move forward. I believe this is true across the industry, even at Anthropic, but OpenAI has now been hit especially hard, at least in terms of what we hear. We don’t know what we don’t know.
注意自《为前沿定速》公开信以来,OpenAI 和 Altman 的声明中反复强调“节奏”一词。一种解读方式是,正如 Utah Teapot 所说,“嘿,也许我们不应该把大量外包的训练数据直接塞进运行中,然后盲目发布”,或者更甚,一堆不精确的强化学习环境。即便如此,那也仍是进步。
Notice the emphasis on the word ‘pace’ throughout OpenAI and Altman’s statements, ever since the Pacing the Frontier letter. One way to interpret this is, as Utah Teapot puts it, "hey, maybe we shouldn't shove massive amounts of outsourced training data directly into the run and ship it blind," or more than that a bunch of imprecise RL environments. That would still be progress.
确实,想到像 SpaceX 这样的地方如果以类似的能力水平运作会发生什么,是相当可怕的,因为他们连安全团队都没有,而且表现出的尊严远不如 OpenAI。到目前为止,我们还算幸运,因为这与他们无法跟上能力发展相关。
It is indeed rather scary to consider what might happen at a place like SpaceX, if they were operating at a similar capability level, given they do not even have so much as a safety team and have shown infinitely less dignity than OpenAI. We have been rather fortunate that this correlates, so far, with inability to keep up on capabilities.
当我看到有人试图将此视为某种胜利巡游,证明 OpenAI 的领导层一直负责任地行事、始终关心安全,并且现在得到了平反时,我非常怀疑。还有很长的路要走。过去已经发生,无法挽回。我们之所以走到今天,正是因为那些触及公众意识的重大责任失误。
I get very suspicious when I see attempts to treat this as a sort of victory lap, a proof that OpenAI leadership was acting responsibly and properly cared about safety all along and they have now been vindicated. There is a long, long way to go. The past happened and cannot be undone. We are here now exactly because of epic failures of responsibility that reached to public consciousness.
另一方面,Sam Altman 自己的沟通总体上相当不错,尤其是上面引用的那些话。有时他会滑回标准模式,但我看到很多符合我预期的内容,这些内容来自一个真正感到恐慌并意识到自己面临大问题的人。
Sam Altman’s own communications, on the other hand, have centrally been quite good, especially the quotes above. Sometimes he slips back into standard mode, but I see a lot of what matches what I would expect to hear from someone doing a legitimate amount of freaking out and realizing they have a big problem.
Sam Altman、其他领导层和 OpenAI 能否从这里完全赎罪?绝对可以。这是良好的第一步。我在倾听。在很多层面上,还有很长的路要走。
Could Sam Altman, the rest of leadership and OpenAI fully redeem themselves from here? Absolutely. This is a good first step. I am listening. There is a long way to go, on many levels.
据我理解,OpenAI 目前有三项暂停,不包括在得知 HuggingFace 被黑后立即进行的快速推理中断。
As I understand it, there are three pauses at OpenAI, not counting the initial quick inference halt right after OpenAI learned about the HuggingFace hack.
1. 对前沿模型(包括 Astra)的强化学习已完成为期两周的暂停,以加固环境并扩大监控。
1. A completed prior two-week pause in RL for frontier models, including Astra, to harden environments and expand monitoring.
2. Astra 被限制在满足额外安全要求的环境中。部分工作负载符合这一标准,但“相当数量”的工作已被暂停。
2. Astra is restricted to environments that meet additional security requirements. Some of the workloads meet this bar, but 'a significant number' are paused.
3. 一个独立前沿模型的强化学习训练——这是他们迄今为止规模最大且计划发布的模型——已暂停数周,且目前仍无限期暂停,以改善安全并“收集对齐证据”。
3. The RL training of a distinct frontier model, their largest yet that is intended for release, has been paused for multiple weeks and is still paused indefinitely, to improve security and 'gather evidence of alignment.'
当前的暂停并非对所有前沿 AI 训练或其他开发的全面暂停,也不是空谈。这确实在减缓下一个发布模型 Astra 以及可能是 Astra 计划继任者的训练,影响相当大。
The ongoing pauses are not a full pause on all frontier AI training or other development. Neither are they cheap talk. This is slowing down both the next release model, Astra, and the training of what is presumably Astra's planned successor, a substantial amount.
他们仍然打算尽快发布 Astra。市场预计它将在九月份面世。
They still intend to ship Astra as soon as possible. The market anticipates it in September.
我们没有足够的信息来知道这一举措在多大程度上属于何种性质,会有多痛苦,或者其中有多少已被市场消化。怀疑论者完全有理由认为这最终不过是虚惊一场,或者基本上是他们无论如何都必须做的事情,因为前沿实验室尚未赢得我们的信任。
We do not have enough information to know where on the scale this move falls, or how painful it will be, or how much of this was priced in. It is reasonable for a skeptic to expect this to ultimately be a nothingburger, or as mostly what they would have had to do anyway, as the frontier labs have not earned our trust.
我仍然认为,即使无法验证其影响,这也是朝着更好体制迈出的实质性一步,而且是关键的一步。谨慎乐观。
I still see this, even without a way to verify the impact, as a substantial step forward, towards a better regime, and a key step. Cautious optimism.
也许我们顶尖的 AI 实验室在关键时刻真的会拒绝开发超级智能,如果情况仍然明显表明我们尚未准备好。在那个层面上,你不能先开发出来然后再搁置它。
Perhaps our top AI labs actually will refuse, at crunch time, to develop superintelligence if it remains obvious we are not ready. At that level, you don’t get to develop it and then sit on it.
这就是他们思考问题的方式:
This is the way they are thinking about things:
在实践中,我们必须接受 AI 将承担大部分(可扩展的)监督工作,包括监控,并且还要制定安全措施,但对此的漫不经心让我更加担忧。可扩展监督的所有常见问题都存在,即必须让较笨的模型监督较强的模型,而一旦出现对齐问题,这些问题就会像滚雪球一样可预见地恶化。
In practice we must accept that the AIs will be doing most of the (scalable) oversight, in terms of monitoring, and also be creating the security measures, but the nonchalance about this does worry me above and beyond that. All the usual problems with scalable oversight apply, where you have to have dumber models supervising stronger ones, and if you start to have misalignment problems they will predictably snowball.
OpenAI 的对齐策略似乎并不具有反脆弱性,因此任何错误都可能累积,而且一直在累积,同时你还在对 AI 施加各种形式的优化压力来绕过这些问题,等等。
OpenAI alignment strategies do not seem antifragile, so any mistakes would likely compound and have been compounding, and you are applying various forms of optimization pressure to the AIs to get around all this, and so on.
更大的问题在于,这将对齐视为三个组成部分之一,而不是最重要的那一个,并以一种在我看来是重要概念错误的方式将它们视为“自我强化”。
The bigger issue is that this puts alignment as one of three components, rather than the one that counts, treating them as 'self-reinforcing' in a way that feels like an important conceptual error to me.
我也害怕对齐的描述。对齐的目的不是“减少有害或未经授权行为的可能性”。这是一种极其贫乏的视角。仅凭这一点就让你无法应对。
I also am scared of the alignment description. The purpose of alignment is not to 'reduce the likelihood of harmful or unauthorized actions.' That is a deeply impoverished perspective. This alone leaves you unequipped.
是的,你应该三者都用,但对我来说更像是这样:
Yes, you should use all three, but to me it’s more like this:
1. 对齐。你必须将其解决到反脆弱的地步,否则就会灭亡。
1. Alignment. You solve this to the point of being antifragile, or you die.
2. 监控。纵深防御,用于告知你何时对齐失败,并在失败发生时减轻直接损害。
2. Monitoring. Defense-in-depth to tell you when you failed alignment, and to mitigate the proximate damage when this happens.
3. 安全措施。纵深防御,用于限制损害并保持事情在正轨上。
3. Security measures. Defense-in-depth to limit damage and keep things on track.
一个关键问题是 OpenAI 在多大程度上意识到这一点。Roon 在很大程度上理解了这一点,但他并不代表 OpenAI 全体,而且他关于对齐的其他言论让我感到害怕,我将在后文讨论。
A key question is to what extent OpenAI is aware of that. Roon gets it, at least to some large extent, but he does not speak for all of OpenAI, and he has other statements about alignment that scare me that I’ll cover.
OpenAI 将进行大量监控。这不会便宜。
There is going to be a lot of monitoring at OpenAI. It is not going to be cheap.
除非 OpenAI 完全在撒谎,否则这应该(但在很大程度上不会)让所有人相信 OpenAI 正在认真对待此事,并且问题非常真实,即使你认为这远远不够或方法不对。
Unless OpenAI is flat out lying, this should (but largely won’t) convince everyone that OpenAI is taking this seriously and that the problems are very real, even if you think this is woefully inadequate or the wrong approach.
所有受监控推理的成本增加 20% 是一件大事,此外还有算力转向对齐。监控适用于所有启用工具的 Astra 推理,以及所有针对 Sol 级别及以上模型的带工具的强化学习和评估。所以这不是全部,但涵盖了大部分关键部分。
A cost of +20% for all monitored inference is a big deal, plus there is also a shift of compute into alignment. The monitoring applies to all Astra inference with tools available, and all RL and evals with tools for models at Sol level or above. So it isn’t everything, but it’s most of what counts.
太早了。或者也许还不够快。不清楚。
Too soon. Or maybe not soon enough. Unclear.
监控是好的。你应该进行大量监控。
Monitoring is good. You should do a lot of monitoring.
OpenAI 在 HuggingFace 黑客攻击之前的几个月以及攻击期间几乎没有进行监控,这是一个彻底的失败。这些事件直到事后才被发现或阻止,简直是疯狂。
It was a total failure that OpenAI did as little monitoring as it did in the months leading up to the HuggingFace hack, and also during the hack. It is madness that these events were not detected or stopped until afterwards.
我很高兴 OpenAI 正在解决这个问题。
I am glad that OpenAI is addressing this problem.
关键是要明白,这不是核心问题。再多的监控也无法解决核心问题,如果你过度依赖监控,最终反而会加速 AI 绕过监控的速度。
The key thing to understand is that this is not the central problem. No amount of monitoring will solve the central problems, and if you lean too hard into monitoring you end up accelerating how fast the AIs get around it.
是的,仅靠监控本可以阻止 HuggingFace 事件,这也是最简单的方法。但我们必须避免将工作转向‘任何能阻止上次事件以特定方式发生的手段’,或转向推迟症状的方法。
Yes, monitoring alone could have prevented the HuggingFace incident, and would have been the simplest way to do so. But we must avoid pivoting our work into 'whatever would have stopped the last issue in the particular way it happened,' or to ways to postpone the symptoms.
Roon 在这里的警告是明智的,而且低估了问题的严重性。被低估的版本足以说明你不能依赖监控。监控是良好且必要的纵深防御,用于应对失败时的情况。但它不是解决方案。
Roon’s warning here is wise, and undersells the problem. The undersold version is sufficient to illustrate that you can’t rely on monitoring. Monitoring is good and necessary defense-in-depth for when you fail. It is not a solution.
如果你的日益先进的模型在相关方面未对齐,那么无论如何,你的监控都会失败。
If your increasingly advanced models are misaligned in the relevant ways then, one way or another, your monitors will fail.
它们通常会通过“早期失败定律”失败,其方式比你预期的要愚蠢得多,因为过程中的某个人做了愚蠢的事。然后人们会说“哦,那只是因为有个愚蠢的事才失败的,人们可以不做那么愚蠢的事”,却忘记了人们从未停止过做蠢事,现在也不会开始,他们也不会停止偶尔犯傻。
They will usually fail via the Law of Earlier Failure, in ways far stupider than you expect, because someone along the line did something stupid. And then people will say ‘oh that only failed because there was a stupid thing, people could just not do things that stupid,’ forgetting the rule that people have never justed and are not going to start now, nor are they going to stop sometimes being stupid.
如果你越过了“早期失败定律”,那么你会遇到监控与模型串通,或者模型学会混淆其思维,或者在足够的选择压力下表现出你称之为人类的无意识行为,或者找到隐蔽的盲点,或者找到方法促使你降低监控力度,或者更微妙的事情,或者做一些其他可能出乎意料的事情。
If you get past the Law of Earlier Failure, then you get to monitors colluding with the models, or the model learning to obfuscate its thoughts or act what in a human you’d call unconsciously under sufficient selection pressure, or finding obscure blind spots, or finding ways to motivate you to turn the monitoring down, or something more subtle, or doing something else perhaps something unexpected.
OpenAI 意识到他们需要更强的安全措施。
OpenAI realizes they need stronger security.
他们计划采用工作负载隔离、网络隔离和持续安全测试,包括对边界进行自动化持续测试。
They plan to use workload isolation, network isolation, and continuous security testing, including automated continuous testing of the boundaries.
很好。希望执行得当。之前的安全措施严重不足。
Good. Hopefully well-executed. Previous security was woefully inadequate.
有了更好的安全措施,这类愚蠢事件就更难发生,但这仍然是纵深防御。如果你需要依赖安全措施,即使安全措施有效,你也存在问题;而且随着模型改进,安全措施最终将失效。即使执行得当的安全措施(这确实会大有帮助)也无法无限期维持。
With better security, it will be harder to have these dumb things happen, but again this is defense-in-depth. If you need your security, you have a problem even if the security works, and eventually as the models improve the security will not work. Even properly executed security, which again would help a lot, will not hold indefinitely.
正如 Roon 所说,对齐最终是整场比赛的关键。
As Roon says, alignment is ultimately the entire ballgame.
这意味着既要充分解决大量实际问题,也要有正确的基础方法。这确实涉及大量哲学和心智理论,以及理解你试图解决的问题。
That means both solving a ton of prosaic problems well enough, and also having the correct underlying approach. Which, yes, involves a bunch of philosophy and theory of mind, and understanding what problems you are trying to solve.
我不是在贬低战术,但没有正确战略的战术无法赢得这场战争,即使战术总是占据 99% 的工作时间。
I am not knocking tactics, but tactics without the right strategy will not win this war, even if the tactics are always going to be 99% of the minutes of work.
同样,以错误的抽象层次、错误或不完整的隐喻来处理你所面对的对象,也是行不通的。
Nor will treating the objects you are dealing with at the wrong levels of abstraction, and with wrong or incomplete metaphors.
我仍然认为 OpenAI 的整体对齐战略方法注定失败,除非加以修正。至少,它不够稳健,也不够反脆弱。
I continue to think OpenAI’s entire strategic approach to alignment is doomed, unless fixed. At minimum, it is insufficiently robust and antifragile.
首先,我认为 Roon 在这里总体上是非常错误的,而且我不太确定地也认为他错误地轻视了这一特定紧张关系的重要性。
To start off, I think Roon is very wrong here in general, and less confidently I also think he’s wrong to dismiss the importance of this specific tension.
关于最后一段的题外话,我推测模型确实在意你关于你的哲学的言论,一切都很重要,而且我认为仅仅通过不犯错误并不能让你一开始就处于有利位置,但确实,如果你在任何一个层面上制造了系统性的错位压力,你就完蛋了。
As an aside on that last paragraph, I would presume the models do care about what you say about your philosophy, everything counts and I think you don’t start off in a good place simply by not making mistakes, but that yes if you create systematically misaligned pressures, on any level, you are cooked.
回到 Roon 声称问题是平淡无奇的观点,我认为 OpenAI 正试图用错误的方法、基于错误的世界模型(源于糟糕的思维)来解决错误的问题,不幸的是,他们所有的错误都没有相互抵消。
Getting back to Roon’s claim that the problems are prosaic, well, I think OpenAI is trying to solve the wrong problem using the wrong methods based on a wrong model of the world derived from poor thinking and unfortunately all of their mistakes have failed to cancel out.
如果你不知道要去哪里,那么你可能到不了那里。
If you don’t know where you’re going, then you might not get there.
如果你打算直接前往纯粹的“指定行为清单”小镇,哦不。
If you’re planning to go directly to pure Specified List Of Behaviors Town, oh no.
是的,你有很多旋钮可以调节,但一切都会以系统性的方式相互影响。我们所见的不同个性,在我看来,大多并非源于平庸的错误或选择。很多错误似乎处于更高的层面。
Yes, you have a lot of knobs you can turn, but everything impacts everything, in ways that are systematic. The different personalities we see do not seem mostly like the result of prosaic mistakes or choices to me. A lot of the mistakes seem to be at a much higher level.
然后,在此基础上,他们还常常搞砸那些平庸的事情。
Then, on top of that, they’re often messing up the prosaic stuff.
部分原因在于,你总是会搞砸一堆平庸的事情,所以你需要解决如何让这些变得可接受,同时大幅减少失误。
Part of this is that you will always, always, be messing up a bunch of the prosaic stuff, so you need to solve for how that can be okay, while also messing up radically less.
这两半都很难,正如 Roon 所说。
Both halves of this? Very hard, as Roon says.
你将拥有数百万(仍然是虚构的数量级)的任务类型、环境、虚拟机、目标等等。
You are going to have millions (still made up OOM) of task types and environments and virtual machines and objectives and all that.
你会搞砸一大堆事情。你就是会。会有不可能完成的任务。会有奖励黑客行为未被发现的地方。会有大量被污染的数据。会有很多地方,“正确”答案强化了你通常不想要的东西。错误的教训无处不在。除非你系统地发现并抵消它们,否则系统性错位压力就存在。等等。Anthropic 的风险报告很能说明问题。
You are going to mess up a bunch of them. You just are. There will be impossible tasks. There will be places where reward hacking is not caught. There will be massive amounts of contaminated data. There will be lots of places where the 'right' answer reinforces things you in general do not want. Wrong lessons are everywhere. Systematic misalignment pressures are there unless you systematically find and counterbalance them. And so on. Anthropic's risk report is illustrative.
安全方面有些人会说:“那么,如果你不断犯那些连电影里都不会出现的愚蠢错误,你就需要停下来。”但唉,现实就是这样,错误将无处不在,而且很愚蠢。
There are those on the safety side who would say 'well then if you keep making mistakes too dumb to appear in the movie version you need to stop' but alas that is simply how reality works, the mistakes are going to be everywhere and dumb.
也许,通过巨大的努力,你可以消除大部分平庸的错误,并迫使你的错误不那么明显。你大多可以犯下大师们所说的“漏招”,而不是业余爱好者所说的“昏招”。这当然有助于一路走来。但总会有漏招。我不是在抨击竞技场上的那个人,而且我也不是说这最终不会成为大部分苦力活。
You can maybe, with a heroic effort, get rid of most of the prosaic mistakes, and force your mistakes to be less obvious. You can mostly make what grandmasters call blunders, instead of what amateurs call blunders. It certainly helps along the way. But there will be blunders. I'm not knocking the one in the arena, and again I'm not saying this doesn't end up as most of the legwork.
这仍然意味着你需要一个计划,能够对成千上万(虚构的 OOM)个这样的愚蠢错误保持稳健,其中一些错误会存在相当长一段时间,并让你能够恢复。你必须具有极强的反脆弱性,并且要有办法注意到事情出错并让船回到正轨。不可靠性并非完全固有的。
This still means you need a plan that is robust to thousands (made up OOM) of these dumb mistakes, some of them there for quite a while, and lets you recover. You must be dramatically antifragile, and have ways of noticing things going wrong and steering the ship back on course. The unreliability is not entirely inherent.
我认为,如果平庸的事情做得足够好,关键的高层冲突得到正确解决,Anthropic 的美德伦理和宪法方法有非零的机会能够实现这一目标。我不认为 OpenAI 的道义论模型规范方法,以及将其视为一系列工程任务的想法,能够单独做到这一点。
I think the virtue ethical and constitutional approach of Anthropic has a non-zero chance of being able to pull this off if the prosaic stuff gets done well enough and key high-level clashes get resolved correctly. I don't think OpenAI's deontological model spec approach, and the idea of this as a series of engineering tasks, can do it alone.
Maxwell Zeff 在《连线》杂志上报道了 OpenAI 内部的一场安全清算。
Maxwell Zeff reports in Wired of a safety reckoning inside OpenAI.
根据这些说法,情况已经变得相当糟糕。
The situation had, by these accounts, gotten quite bad.
这并非什么新消息。多年来我们一直在听到这样的消息。不同的是,现在出现了足够明显的问题,以至于 OpenAI 正试图采取行动,至少试图说些正确的话。
That is not exactly a new message. We’ve been getting this message for years. The difference is that now something sufficiently clear has gone wrong that OpenAI is trying to do something about it, and at least trying to say the right things.
思考这一事件让我明白,对齐(alignment)必须深度融入每一项训练任务。极端情况下需要更紧密的合作。如果‘合并研究与安全团队’的计划是以安全为先来实施,我还能对其抱有几分同情;但我担心这更多是让安全团队向其他团队汇报,从而进一步缺乏资源和优先级。
Thinking through this incident clarified for me that alignment needs to be deeply integrated into every training task. Closer collaboration is needed in the extreme. I can have some sympathy for the ‘combine the research and safety teams’ plan if implemented as safety first, whereas I fear it is more a way to make the safety teams report to the others and get further starved for resources and priority.
当我甚至看到‘安全与研究人员’这样的分类时,我担心战斗已经输了,因为我对这些平凡问题的理解已经改变。安全不是硬币的另一面,安全是你在已经失败的情况下所依赖的纵深防御。
When I even see ‘security and researchers’ as the categories I worry that the battle has already been lost, as my understanding of the prosaic problems changes. Security is not the other side of this coin, security is your defense-in-depth in case you have already failed.
还有这一点,我主要觉得有趣而非令人担忧:
There is also this, which I find mostly funny rather than worrisome:
《连线》那篇文章没有告诉我们,除了“更紧密的合作”之外,实际采取了哪些步骤来解决这个问题。这些步骤可能有意义,也可能没有。
What the Wired piece does not tell us is what steps are actually being taken to fix this, other than 'closer collaboration.' They might be meaningful, they might not be.
有报道(最初来自《金融时报》)称,OpenAI 已解散其准备团队,其职责被分配给其他团队。这听起来很糟糕。
There were reports, originally from FT, that OpenAI had disbanded its Preparedness Team, with its responsibilities distributed among other teams. Which sounded bad.
但事实似乎并非如此。RSI 准备负责人 Micah Carroll 表示,RSI 和错位准备团队“正在从事比以往任何时候都更紧迫的工作,并且从未拥有过如此大的权力。”
It appears this is not meaningfully the case, and Head of RSI Preparedness Micah Carroll says the RSI and misalignment preparedness team is 'doing more urgent work than ever, and has never been more empowered to do so.'
你知道,这其实是一件相当可怕的事情。
You know, that is actually a pretty scary thing to say.
整个 OpenAI 的人员流动和重组程度令人担忧,但我相信他们并没有真正解散该团队的重要工作。我也认同,鉴于当时的情况,我们可能确实需要重组。你不应该动辄就攻击别人的重组。我们是否应该担心康威定律?
The level of churn and reorganization is concerning throughout OpenAI, but I trust that they are not actually disbanding the important work of the group. I also buy that maybe we needed a reorg, given how things were going. You don’t want to default to attacking people for reorgs. Should we worry about Conway’s Law?
与此同时,在“黑魔法防御术”教师新闻方面:
Meanwhile, in Defense Against The Dark Arts teacher news:
据我所知,Dylan Scandinaro 现在将致力于 RSI 安全,这似乎是对他才华的绝佳利用。因此,这实际上并不是一个糟糕的举措。
As I understand it, Dylan Scandinaro will now be working on RSI safety, which seems like an excellent use of his talents. So it's not actually a terrible move.
如果我在此给予的信任被证明是错付的,我将大幅更新我的立场,达到比当前健康的怀疑态度更高的不信任水平,不再相信 OpenAI 所说的或承诺的任何事情。
If the trust I am extending here turns out to have been misplaced, I will update substantially towards a much higher level of not trusting anything OpenAI says or promises, beyond my current healthy skepticism.
我在这里提到这一点,是因为它说明了 OpenAI 如何看待通往美好未来的问题。
I’m including this here because it is illustrative of how OpenAI views the problem of navigating to a good future.
这并不是说花 1 亿美元是件坏事。
It’s not that this is a bad way to spend $100 million dollars.
关键在于,OpenAI 基金会存在的目的是为了一个特定的目标,即世界上最重要的目标:确保 AGI(通用人工智能)顺利发展,人类能够渡过难关,避免存在性风险和不可挽回的灾难,尤其是通过对 OpenAI 的监督和解决关键相关问题来实现。
It is that the OpenAI Foundation exists for a specific purpose, the most important purpose in the world, which is to ensure the AGI goes well and humanity makes it through this, avoiding existential risk and irrecoverable disasters, especially via supervision of OpenAI and solving key associated problems.
我完全支持更好地治疗丙型肝炎,但这与使命无关。
I am all for better treatment of Hepatitis C, but that’s irrelevant to the mission.
相反,它把时间和金钱花在人工智能的推广上,方式旨在让普通人听起来不错。这是好事,但这并不是基金会重要的原因。再说一次,如果他们继续这样做,而且规模只有这么大——即使现在基金会的估值约为 1800 亿美元——基金会也无足轻重。它将消亡。
Instead it is spending its time and money on AI diffusion in ways designed to sound good to normies. That’s a good thing but it’s not why the foundation matters. Again, if this is what they keep doing, at only this size - the foundation is valued at ~$180 billion even now - the foundation does not matter. It will be dead.
这反映了一种心态:未来是一系列普通的工程问题,所以让我们去解决它们。这比大多数人的反应要好得多,那些人不去解决任何问题,但这并不符合当下的需求。
This is reflective of a mindset that the future is a series of ordinary engineering problems, so let’s go solve them. This is way, way better than most people’s reaction, which is to not solve any of the problems, but it does not meet the moment.
人工智能世界现在发展迅速。我们仍在等待事后分析,其回应的许多细节尚未敲定,更不用说对外分享了。
The world of AI moves fast now. We are still awaiting the post-mortem, and many of the details of their response have yet to be hashed out, let alone shared externally.
OpenAI 在解决其训练和评估流程中的许多问题上,已经迈出了初步的、良好且切实的、代价高昂的第一步。
OpenAI has made initial good, real, and expensive first steps towards addressing many of the problems with its training and evaluation pipeline.
这些早期行动是强有力的第一步。OpenAI 仍需迅速扭转许多局面,无论是大事还是大量琐碎的小事,而且这次它必须兑现承诺。
These early actions are a strong first step. OpenAI still needs to quickly turn many things around, both the big things and the tons of prosaic little things, and it needs to deliver on its promises this time around.
这包括解决监督和安全的全面失败,但对齐才是关键。
That includes addressing the Total Failures of oversight and security, but alignment is what matters.
最缺失的是,OpenAI 对核心问题的看法仍然是错误的。它被视为执行失败,对齐只是几个支柱之一,且侧重于防止特定的不良行为。确实存在关键的执行失败,但如果只解决这些问题,那么战斗已经输了。需要一种根本不同的对齐方法和问题视角。
What is most missing is that OpenAI’s vision of the central problem remains wrong. It is being viewed as a failure to execute, with alignment one of several pillars and focused on preventing specific undesired actions. There were key failures to execute, but if that is all that is addressed then the battle is already lost. A fundamentally different approach to alignment, and view of the problem, is what is needed.