Anthropic Risk Report: August 2026
打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→这份 2026 年 8 月的 Anthropic 风险报告评估了高级 AI 模型的风险,重点关注自主性、自动化 AI 研发以及生物/化学武器生产,但不包括网络风险。报告详述了两个内部模型,其中模型 2 能力更强,但因潜在危险或专门化而未发布。报告通过一致性、普遍性和自然涌现等维度定义了错位,并引入了高风险场景的威胁模型。关键关注点包括故意与非故意伤害,如近期 OpenAI 的失败所示,以及即使在不便时也需遵守的稳健风险阈值。作者批评了报告对特定威胁模型的狭隘关注,主张更广泛地考虑非故意伤害以及安全部门的重要性。总体而言,报告提供了评估 AI 风险的框架,但强调了在能力开发与安全之间平衡的持续挑战,敦促对规则变更设定更高门槛,并将安全更好地整合到所有 AI 研究中。
This August 2026 Anthropic risk report evaluates risks from advanced AI models, focusing on autonomy, automated AI R&D, and biological/chemical weapons production, while excluding cyber risks. It details two internal models, with Model 2 being more capable but not released due to potential dangers or specialization. The report defines misalignment through dimensions like coherence, pervasiveness, and natural emergence, and introduces threat models for high-stakes settings. Key concerns include intentional versus incidental harm, as seen in recent OpenAI failures, and the need for robust risk thresholds that are honored even when inconvenient. The author critiques the report's narrow focus on specific threat models, arguing for broader consideration of incidental harms and the importance of safety departments. Overall, the report provides a framework for assessing AI risks but highlights ongoing challenges in balancing capability development with safety, urging higher bars for rule changes and better integration of safety across all AI research.
3. 规则是严肃的,但不是字面上的。
3. The Rules Are Serious But Not Literal.
4. 错位是一种心态(2.5)。
4. Misalignment Is a State of Mind (2.5).
5. 自主性威胁模型 1:高风险环境中的错位(2)。
5. Autonomy Threat Model 1: Misalignment in High-Stakes Settings (2).
6. 我以前不知道的“安全”一词的某些奇怪用法。
6. Some Strange Uses Of The Word Safe I Wasn’t Previously Aware Of.
9. 第 2 节中其余的重要论点。
9. The Rest of the Important Arguments In Section 2.
11. 内部部署前审查(2.18)。
11. Pre-Internal-Deployment Review (2.18).
12. 内部使用监控指南(2.23.1)。
12. A Guide To Internal Use Monitoring (2.23.1).
14. 权力寻求环境评估(2.24)。
14. The Power Seeking Environment Evaluation (2.24).
16. 自主性威胁模型 2:自动化研发的风险(3)。
16. Autonomy threat model 2: Risks from automated R&D (3).
19. 我们可能在多大程度上加速我们的 AI 研究人员?
19. How Much Could We Be Accelerating Our AI Researchers?
20. 如果我们替换掉研究人员,可能会出什么问题?
20. What Could Possibly Go Wrong If We Replaced Our Researchers?
21. 人工智能研发自动化的风险缓解措施。
21. Risk Mitigations for AI R&D Automation.
22. 人工智能研发自动化的总体风险。
22. Overall Risk from Automation of AI R&D.
23. 生物武器以及技术上同样适用的化学武器生产。
23. Biological and Technically Also Chemical Weapons Production.
24. 生物和化学武器的威胁模型。
24. The Threat Models for Biological and Chemical Weapons.
30. 拒绝寻找创新的错位技术(5.2.2)。
30. Refusing To Find Innovative Misalignment Techniques (5.2.2).
31. 将思维链推理大量暴露于评分压力之下(5.2.3)。
31. Exposing the Chain of Thought Reasoning To Grading Pressure Quite a Lot (5.2.3).
32. 在生产训练运行期间直接训练错位行为(5.2.4)。
32. Directly Training On Misaligned Behavior During a Production Training Run (5.2.4).
33. 一个能够访问敏感资源的无监控、不受限制的智能体实例(5.2.5)。
33. An instance of unmonitored unrestricted agents with access to sensitive resources (5.2.5).
34. 在错位伪造转录数据集上重复训练(5.2.6)。
34. Repeated training on alignment-faking transcript datasets (5.2.6).
35. 从 Anthropic 作为前沿 AI 公司的运营中获益(5.3)。
35. Benefits From Anthropic’s Operating as a Frontier AI Company (5.3).
本报告中提到了两个内部模型。
There are two internal models referenced in this report.
模型 1 与 Mythos Preview 和 Mythos 5 大致相似,预计不会在外部或内部广泛部署。
Model 1 is broadly similar to Mythos Preview and Mythos 5, and is not expected to have wide deployment, externally or internally.
模型 2“比 Mythos 5 能力稍强”,在内部使用上有明显改进,但“并未展现出从 Claude Opus 4.6 到 Mythos Preview 所观察到的能力跃升程度”。该模型将仅供内部使用。
Model 2 is 'somewhat more capable than Mythos 5,' a noticeable improvement for internal use, but 'does not display a capability jump of the degree observed from Claude Opus 4.6 to Mythos Preview.' This model is going to be internal-use only.
从 Opus 4.6 到 Mythos Preview 的跃升是巨大的,跨越了多个发布周期。报告中我们获得了更多数据,但鉴于模型 2 被拿来与这一跃升作对比,可以推测在内部用途上,模型 2 领先了多个版本。
The jump from Opus 4.6 to Mythos Preview was big. Several release cycles big. We get a lot more data throughout the report, but given Model 2 gets that comparison point, one would presume that for internal purposes Model 2 is multiple releases ahead.
与此形成对比的是,在 AECI 上它仅领先 1.5 分,这大约只相当于一个月的进展。我猜测这是一个低估,而且模型 2 出于其目的,并未刻意配备大量通用技能。而在替代 Anthropic 研究人员的测试中,我们看到从 54.8%(Mythos Preview)和 50.3%(Mythos 5)大幅跃升至模型 2 的 62.8%。
To contrast with that, on the AECI it is only 1.5 points ahead, which would be only about a month of progress. My guess is this is a low estimate, and that Model 2 was intentionally not decked out with a bunch of general skills given its purpose. Whereas on the test of substituting for Anthropic's researchers we see a big jump from 54.8% (Mythos Preview) and 50.3% (Mythos 5) to 62.8% for Model 2.
不发布模型 2 存在多个合理的充分理由。模型 2 可能高度专门用于内部任务。它也可能能力过强或过于危险而不宜发布,无论是出于危害的考虑,还是担心它会加速他人的进展。Fable 5 相对较低的采用率表明,人们倾向于保留这样的模型。
There are multiple plausible good reasons for not releasing Model 2. Model 2 could be highly specialized for internal tasks. It could also be too capable or dangerous to release, either due to harm or the worry it would accelerate others. The relatively low adoption rate for Fable 5 points towards wanting to hold such a model back.
该报告按风险类型、威胁模型、相关 AI 模型进行划分,并详细说明了当前的能力、行为、缓解措施和总体风险水平。然后展望未来,并提出全行业建议。
The report is divided by each type of risk, by threat model, by relevant AI models, and details current capabilities, behaviors, mitigations and overall level of risk. It then looks forward to the future and offers industry-wide recommendations.
风险报告涵盖截至一个月前的事件,因此覆盖日期为 2026 年 7 月 15 日。一年前这似乎还可以。现在看起来已经相当久了,是吧?
Risk reports cover events up to a month prior, so the coverage date is July 15, 2026. A year ago that would have seemed fine. Now it seems like quite a while, huh?
长期利益信托(TLBT)有权要求对报告进行外部审查,但尚未这样做。如果我是他们,我至少下次会这样做。
The Long-Term Benefit Trust (TLBT) is authorized to request external review of the report, but has not done so. I would at minimum do so next time if I was them.
我喜欢他们既涵盖“这带来的绝对风险是什么”,也涵盖“鉴于其他所有人的存在,这带来的边际风险是什么”。
I like that they cover both 'what is the absolute risk from this' and 'what is the marginal risk from this given everyone else exists.'
他们将自主性、自动化 AI 研发以及生物和化学武器生产视为风险。即使到现在,将网络风险排除在核心威胁模型之外仍然很奇怪。我理解其本身不足以构成灾难的论点,但在实践中,我认为你需要处理它。
They deal with autonomy, automated AI R&D and biological and chemical weapons production as risks. It is odd, even now, to exclude cyber risks from the core threat models here. I understand the argument from insufficiently catastrophic on its own, but in practice I think you need to deal with it.
在多个地方,风险报告实质上承认了我反对意见的某些版本,但随后又忘记了这些承认,并未改变其结论。
At several points, the risk report essentially concedes versions of my objections, but then forgets that it conceded them and doesn’t alter its conclusions.
然后在第 5 节中,他们转向了更有趣的问题。
Then in Section 5 they get to the more interesting questions.
风险阈值的定义已从较简单的表述演变为包含更多具体细节的表述。
The risk threshold definitions have changed from simpler statements to ones with more specific detail.
对于 AI 研发自动化,我认为新措辞更难解析,但尚可。对于新型生物武器生产,新版本将范围缩小到仅关注“替代稀缺的人类专业知识”,而排除了其他方法。这是一个合理的首要威胁模型或场景,但我担心只关注这一条路径。
For AI R&D automation, I think the new wording is harder to parse but fine. For novel biological weapons production, the new version narrows to only look at 'substitute for the scarce human expertise' to the exclusion of other methods. This is a reasonable primary threat model or scenario, but I worry about looking only at that path.
好消息和坏消息并存的是:我认真对待此类文件,但我不再字面理解此类文件,无论朝哪个方向。
The both good and bad news is I take such documents seriously, but I no longer take such documents literally, in either direction.
1. 如果 Anthropic 或其他前沿实验室的模型在技术上超过了其风险阈值,但方式看似无害,我预计他们会修改阈值或找到其他方式忽略这一事件。
1. If Anthropic or another frontier lab has a model that technically crosses their risk threshold, but in a way that seems to be harmless, I expect them to modify their threshold or find some other way to ignore this event.
2. 如果 Anthropic 或其他前沿实验室的模型在技术上未超过其风险阈值,但显然在规则试图衡量的方面存在风险,我预计他们会像超过阈值一样采取行动。
2. If Anthropic or another frontier lab has a model that technically does not cross their risk threshold, but is obviously risky in the way the rules tried to measure, I expect them to act as if it had crossed the threshold.
一如既往:我希望即使在你认为“不应该算”的情况下,阈值也能得到尊重,因为这就是承诺的意义所在。而且你应该预料到,在匆忙发布时会出现各种合理化解释,你需要在这方面树立一个好榜样。最初的想法是“如果-那么”承诺,即事先约定:如果发生[X],你就做[Y],但我们的文明似乎缺乏这种技术。我们不能,我们不知道怎么做。
As always: I would like to see thresholds honored even in cases where you think it 'should not count,' because that is what commitments are for, and you should expect rationalizations to occur as you rush to make releases, and you need to set a good example on this. The whole original idea was if-then commitments, where if [X] happens you do [Y], agreed upon in advance, but our civilization seems to lack this technology. We can't, we don't know how.
我并不是说你们永远不会修改规则使其更宽松,如果你们有充分证据表明应该这样做的话。如果规则永远不能变得更宽松,那么规则就必须永远不严格。我们不希望那样。但是,改变规则的门槛,尤其是在应对即将发生的事件时,需要比现在高得多,而且我们需要愿意承担一些实际代价,即使我们现在认为这很愚蠢。
I'm not saying you would never modify your rules to be more lenient, if you have good evidence that you should do that. If you can never make the rules more lenient then the rules have to never be strict. We don't want that. But the bar for changing the rules, especially in response to a pending event, needs to be a lot higher than it is, and we need to be willing to endure some actual costs even when we now think it is dumb.
目前,Anthropic 和 OpenAI 在应对威胁以及根据当时已知信息决定何时发布什么、不发布什么方面,似乎做出了不错的决策,即使他们在其他方面犯了错误。这很好。我希望这种情况能持续下去。
For now, Anthropic and also OpenAI have made what seem like good decisions in terms of responding to threats and deciding what to release and not release when, based on what is known at the time, even if they make mistakes elsewhere. That is good. I hope it continues.
定义部分非常值得赞赏。
The definition section is greatly appreciated.
我非常喜欢这个定义,尽管我有一些重要的吹毛求疵之处:
I like this definition a lot, although I have important nitpicks:
我的第一个修改是将“宪法”泛化,以涵盖模型规格和其他系统的预期属性,这样它就可以应用于 Anthropic 之外。
My first change would be to generalize 'constitution' to include model specs and other intended properties of the system, so this can be applied outside Anthropic.
我的第二个修改是明确要求理性人经过反思后认为该计算因此是不可接受且不合理的。也就是说,在某些情况下,如果有充分的理由超越这些,那么违反法律或做某些“令人反感”的事情或违反模型的宪法并不明显是错位的。在某些情况下,所有路径都会这样做。
My second change would be to specify that this requires the reasonable person to find the computation therefore, on reflection, unacceptable and unjustified. As in, it is not obviously misaligned to break the law or do something 'objectionable' or violate the model's constitution, under some circumstances, if there are good reasons that override this. In some circumstances, all paths will do this.
在这种观点下,错位不是“诚实的错误”或能力失败或缺乏足够信息。
Misalignment, in this view, is not an 'honest mistake' or failure of capability or lack of sufficient information.
正如 Nate Soares 在推特上指出的,根据这一定义,一个 AI 可能并非最初就错位,但依然会造成伤害并导致错位状态,例如,AI 在启动一堆子智能体(或训练新模型等)时草率行事。这里的目标是将其视为一种能力或执行失败,导致后续的错位失败,但这本身并非错位。
As Nate Soares points out on Twitter, you can do harm and lead to a misaligned state without being originally misaligned per this definition, such as an AI being sloppy about spinning up a bunch of subagents (or training a new model, etc). The goal here is to think of that as a capabilities or execution failure that leads to a subsequent misalignment failure, but this itself is not misalignment.
或者,正如其他人所说,永远不要把可归因于无能的归因于恶意。但同样,无能往往导致恶意。
Or, as others have said, never attribute to malice what can be attributed to incompetence. But also incompetence often is what leads to malice.
其他定义对此进行了细分,从而说明了思考(我的转述):
Other definitions subdivide this, thus illustrating thinking (my paraphrases):
1. 一致性。如果你与某个我们不喜欢的目标或偏好集的实现保持一致,那么你就是“一致性错位”。否则,你就是非一致性错位。
1. Coherence. You are ‘coherently misaligned’ if you are aligned to the achievement of some goal or set of preferences that we disfavor. Otherwise you are incoherently misaligned.
1. 鉴于人类的例子,我们不应期望当前能力水平下的任何心智完全一致,因此这大概是我们应用于相对“一致”的人类战略行动的那种宽松标准。
1. Given the example of humans, we should not expect any minds at current capability levels to be fully coherent, so presumably this is to the kind of loose standard we would apply to a relatively ‘coherent’ human acting strategically.
2. 普遍性。普遍性错位发生在广泛的输入范围内,尽管#并非所有输入,我们期望它在训练或评估期间显现。足够狭窄的必要条件,以至于在那段时间内可能被忽略,称为上下文相关。
2. Pervasiveness. Pervasive misalignment occurs over a broad range of inputs, although #NotAllInputs, and we would expect it to show up during training or evaluation. Narrow enough required conditions to be plausibly missed during that time is called context-dependent.
1. 明显的问题是,上下文可能是“你确信自己不在训练中,也不在评估中”。
1. The obvious problem is the context could be 'you are confident you are not in training and not in an eval.'
2. 或者甚至“你有能力从中获得很多好处”。
2. Or even 'you have the leverage to gain a lot from doing this.'
3. 因此,这似乎并不是一种明显的、在关节处划分现实的好方法。
3. Thus, it does not seem obviously like a good way to divide reality at its joints.
4. 这似乎也不是一种让你对自己的处境感觉好很多的好方法。
4. Nor does it seem like a good way to feel much better about your situation.
3. 自然涌现。如果没有任何心智(AI 或人类)有意为之,错位就是自然涌现的。如果是有意为之,包括通过数据投毒,则是被设计的。
3. Naturally emerging. Misalignment is naturally emerging if no mind (AI or human) did it on purpose. It is engineered if it was done on purpose, including via data poisoning.
1. 在某种程度上,许多人和团体一直在以各种方式试图对人和 AI 进行‘数据投毒’,这构成了互联网活动的很大一部分。
1. To some extent, many people and groups are constantly trying to 'data poison' both humans and AIs in various ways, as this constitutes a large percentage of internet activity.
2. 这仍然感觉像是一种‘看到就知道’的色情式情况,我们指的是专门设计来破坏 AI 的数据,或是另一种故意将恶意内容植入模型的尝试。普通的‘用支持某产品的言论淹没互联网’的阴谋不算。
2. This still feels like a pornography-style 'I know it when I see it' situation, where we mean data specifically designed to disrupt the AI, or another deliberate attempt to bake something malicious into the model. Ordinary 'flood the internet with pro-widget statements' plots don't count.
4. 已知。如果你有理由认为某种形式和程度的错位可能存在,那么它就是已知的;否则就是未知的。
4. Known. A form and magnitude of misalignment is known if you have reason to think it likely exists; otherwise, it is unknown.
1. 我有点希望这也能考虑‘惊讶’的程度?
1. I kind of want this to also consider the amount of 'surprise'?
2. 也就是说,你不知道但可能存在的事物,与你认为不存在或极不可能存在的事物之间,存在巨大差异。
2. As in, there is a big difference between things you don’t know about but that could exist, and things you think you know do not exist, or that you think are highly unlikely to exist.
5. 严重性。可能促成优先风险路径。
5. Severe. Could contribute to a priority risk pathway.
6. 致害性。造成预期伤害,无论伤害是否有意。
6. Harm-Inducing. Causes expected harm, whether or not the harm is intended.
7. 高风险性。促成优先威胁路径。
7. High-Stakes. Contributes to a priority threat pathway.
1. 这似乎应该更一般化,严重性也应如此,我担心否则会导致混淆。虽然我理解这一点。
1. This seems like it should be more general, as should severe, and I worry that this is going to otherwise cause a conflation. Although I do get it.
8. 未缓解。指那些未被预防、从而造成预期伤害的潜在有害影响。
8. Unmitigated. The potentially harm-inducing effects that have not been prevented, and that thus cause expected harm.
1. 我想说这应该是风险中未缓解的部分?
1. I want to say this should be the portion of the risk that is unmitigated?
好的定义很难。这些是扎实的初步尝试。一个持续的担忧是,它们将所选示例威胁具体化为真正重要的威胁,而不是某些重要的威胁。
Good definitions are hard. These are solid first attempts. The consistent worry is that they reify the idea that the chosen example threats are the threats that matter, rather than some of the threats that matter.
这涉及主动错位的输出。
This is about actively misaligned outputs.
基本上:自主性 1 是指 AI 自主地、故意地造成永久性损害。
Basically: Autonomy 1 is AI permanently damaging things on its own, on purpose.
该威胁模型聚焦于故意伤害,而非 AI 追求自身目标过程中附带造成的伤害。
The threat model is focused on intentional harm, rather than the AI going after its own ends in ways that require incidental harm.
该场景明确排除了 AI 编写草率代码或犯“诚实错误”的情况。
The scenario explicitly excludes AI writing sloppy code or making ‘honest mistakes.’
因此,在没有“持续错位”的情况下,你是安全的这一结论是成立的。
Thus the conclusion that you are safe absent 'persistent misalignment.'
最近 OpenAI 的全面失败最终导致了 HuggingFace 被黑客攻击,这说明了需要审视更广泛的版本,包括附带伤害。黑客攻击本身可以称为“故意伤害”,但我们看到的主要是附带伤害,包括训练管道的附带损坏,以及随之而来的事实上的“安全研究”。
The need to look at the broader version, involving incidental harm, has been illustrated by the recent total failures at OpenAI that ultimately led up to the HuggingFace hack. The hack itself could be called 'intentional harm' but mostly what we saw was incidental harm, including an incidental corruption of the training pipeline, and with it de facto 'safety research.'
理论上你可以“改变安全研究的结果”,但如我们所见,以其他方式改变训练管道内的情况可能产生相同或更糟的后果。能力和安全工作的区分很可能正是导致这些问题的原因之一,而且我们通常不需要这种特定的针对性。
You could in theory 'alter the results of safety research' but altering the situation within the training pipeline in other ways can have the same effects or worse, as we saw. The differentiation of capabilities and safety work within that is plausibly a lot of what is causing these problems, and generally we do not need this specific targeting.
在我看来,这非常符合“高风险场景中的错位”。
That to me very much was 'misalignment in a high-stakes setting.'
埃隆·马斯克几个月前表示,xAI 没有安全部门,因为每个人都对安全负责。这相当疯狂,尤其是如果这是真的。你需要一个安全部门,但你不能认为能力人员不也大部分属于安全部门。
Elon Musk said some months ago that xAI does not have a safety department because everyone is responsible for safety. That is rather bonkers crazy, especially if true. You need a safety department, but you can’t think of the capabilities people as not also largely in the safety department.
具体的威胁模型和关注点是有帮助的。当你挑选出三个威胁模型时,担心的是你会这样做而排除其他威胁模型。
Concrete threat models and focus points are helpful. The worry when you pick out three threat models is that you will do this to the exclusion of other threat models.
他们提供了八条具体路径,这是我们之前见过的列表:
They offer eight specific pathways, a list we’ve seen before:
与 OpenAI 的事件相比,其中许多在所需的邪恶程度上感觉有些‘卡通化’,但历史上现实世界中的顶级恶棍确实常常表现得相当卡通化,只不过因为具体行动是真实的,我们往往不再认为它是卡通化的。
Something about many of these feels ‘cartoonish’ in its required villainy, when compared to the events at OpenAI, but the top real world villains have indeed often acted quite cartoonishly throughout history, except that because the specific action is real we often stop thinking of it as cartoonish.
不完全是,但基本上这些论点是在解释:尽管已知当前模型存在错位(misalignment)的方式,而且可能还有其他未知的错位方式,但模型可能尚未在特定方面‘普遍’错位,也不具备能够实现这些路径的能力。模型缺乏强大的‘隐蔽能力’(covert capabilities)。
Not quite, but basically the arguments are why, despite the known ways current models are misaligned, and there probably being additional other ways they are not aligned, the models are probably not yet 'pervasively' misaligned and capable in the particular ways that enable these pathways. The models lack strong 'covert capabilities.'
我认为‘上下文相关’的错位并不比非上下文相关的错位危险性低多少,甚至可能更危险。Anthropic 似乎认为这种差异非常重要,因此才有了这样的图表,尽管他们在定义中也承认这种差异并未明显起到实际作用。
I don't think 'context-dependent' misalignment is that much less dangerous than it being not context dependent. It might even be more dangerous. Anthropic seemingly thinks this difference matters a lot, hence charts like this, despite conceding in the definitions that this difference does not obviously do any real work.
[](https://substackcdn.com/image/fetch/$s_!KtAd!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5b4f82bd-eaca-4dde-916d-e05aa9779d74_970x598.png)
[](https://substackcdn.com/image/fetch/$s_!KtAd!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5b4f82bd-eaca-4dde-916d-e05aa9779d74_970x598.png)
或者,‘不像其他一些实验室的白痴,我们偶尔会对记录的对话进行自动化监控,所以如果模型发疯,我们,你懂的,会注意到。’
Or, 'unlike those idiots at some other labs, we occasionally do automated monitoring of our logged transcripts, so if the models were going nuts we would, y'know, notice.'
我的回应是:当你们运行一次评估时,AI 有 141,006 次意外开放了互联网访问权限,其中包括三次入侵真实网站的事件,而你们直到 OpenAI 事件引发大规模回顾性审查时才注意到。
My response is that you did not notice when you ran an eval where 141,006 times the AI had open unintended access to the internet, including three incidents of hacking real websites, and did not notice until there was a large-scale retrospective review triggered by events at OpenAI.
Anthropic 明确指出,该风险报告仅涵盖当前模型的对齐失败风险,尽管这包括神秘且更先进的模型 2。
Anthropic is clear that the risk report covers misalignment risks only from present models, although this includes the mysterious and more advanced Model 2.
报告确实包含了当前对齐失败可能导致未来模型对齐失败的方式。
It does include ways misalignment now could cause future models to be misaligned.
我认为这样的风险报告最好具有前瞻性,提供对未来模型风险的评估,以及需要采取的措施,与我们有能力应对的方面进行对比。但长报告就是长,我不想让完美成为优秀的敌人,所以我大体上尊重这一选择。
I think it would be good for a risk report like this to be projecting, and offer assessments of where they see future risk from future models, and what needs to be done about it versus where we are on pace to be fine. But Long Report Is Long and I don’t want perfect to be enemy of the good, so I’ll mostly respect the choice.
1. 风险等于未缓解的、由错位计算造成的总预期(EV)灾难性危害。
1. Risk equals total expected (EV) unmitigated catastrophic harm from misaligned computations.
1. 这是一个会计恒等式。我同意它是正确的。但尚不清楚这是否是一种有用的思考方式。
1. This is an accounting identity. I agree it is true. It is not obvious it is a useful way to think about the situation.
2. 也就是说,是的,你可以在抽象层面将各个计算的预期危害相加,但我不知道如何做到这一点,也不知道如何使用它。
2. As in, yes, you can in the abstract sum the expected harm from the individual computations but I don’t know how to do that, or how to use it.
2. 主张 1:模型不太可能具有强大的隐蔽能力。
2. Claim 1: Models are unlikely to have strong covert capabilities.
1. 他们没有定义“隐蔽”,但按照常识定义,我认为目前这可能是正确的,尽管我对所呈现的具体证据的价值持怀疑态度。随着时间的推移,这一论断将变得不那么正确,且概率更低。
1. They don’t define ‘covert’ but by a common sense definition I think this is probably true for now, despite my skepticism of the value of the particular evidence presented. It will get less true, with lower probability, over time.
2. 证据表明,模型在不透明推理方面表现不佳,更不用说思维链(CoT)中的内容、动作混淆和选择性动作了。
2. The evidence is that the models are poor at opaque reasoning, not mentioning things in chain-of-thought (CoT), action obfuscation, and selective action.
3. 在评估中,这当然就是新的“小鼠实验”,因为如果存在强大的隐蔽能力,可能包括评估意识,尽管 NLA 证据提供了一些安慰。
3. In evals, which of course is the new in mice, because if there were strong covert capabilities that could include eval awareness, although the NLA evidence provides some reassurance.
4. 这种设置还假定你主要通过推理轨迹来检测错位,这本身就有令人担忧的含义,包括如果你不监控思维链会发生什么,或者模型能否推断出其思维链何时未被监控,以及你开始与训练模型混淆的行为共舞。
4. This setup also presumes you are detecting misalignment primarily via reasoning traces, which has its own scary implications, including what happens if you are not monitoring the chain of thought, or if the model can infer when its chain of thought is not monitored, and that you are starting to dance with training the models to obfuscate.
5. 他们在 2.16.1 节中指出,他们担心自己对隐蔽能力水平的判断有误。这很合理,因为能力本身是隐蔽的。
5. They note in 2.16.1 that they are concerned they are wrong about the level of covert capabilities. Which makes sense, since the capabilities would be covert.
6. 他们在 2.21 节中指出,如果主张 1 为假,且模型具有强大的隐蔽能力,那么许多其他主张也将因此失效。我同意这一点,也同意相反的情况:如果我们能相信主张 1 是强大且广泛成立的,对于隐蔽的广义定义,那将在许多方面有所帮助。
6. They note in 2.21 that if Claim 1 is false, and the models had strong covert capabilities, that many other claims would therefore fail. I agree, and would also agree to the reverse, that if we could believe in a strong and strongly general Claim 1, for a broad definition of covert, that would help in many areas.
7. 然而,即使完全没有隐蔽能力,也不会给我带来高度的保证,因为我认为还有其他途径可以绕过这一点。
7. Even a complete lack of covert capability would not, however, provide me with high assurance, as I believe there are other pathways around this.
3. 覆盖风险可以分解为已知错位、未知严重普遍错位和未知严重情境依赖错位所带来的风险。
3. Covered risk can be decomposed into risks from known misalignment, unknown severe pervasive misalignment, and unknown severe context-dependent misalignment.
1. 这是另一个会计恒等式,因为根据定义,每个风险必须恰好是这三者之一,因为它要么是[普遍或情境依赖],要么是[已知或未知]。
1. This is another accounting identity, in that by definition every risk must be exactly one of these three things, since it is either [pervasive or context-dependent] and either [known or unknown].
4. 每个覆盖风险项可以分解为发生概率、预期影响和其他因素。
4. Each covered risk term can be factored into probability of occurrence, expected impact, and other factors.
7. Even full lack of covert capability would not however provide me high assurance levels, as I think there are other pathways around this.
给定发生情况下的危害,以及未缓解危害的比例。
Harm given occurrence, and fraction of harm unmitigated.
1. 再说一遍,会计恒等式,当然,前提是你能衡量这些。
1. Again, accounting identity, sure, if you could measure these things.
5. 主张 2:已知错位的预期危害较低。
5. Claim 2: Expected harm from known misalignment is low.
1. 小心,伊卡洛斯,但如果我们仅限定于主动的灾难性危害,这是合理的。
1. Careful, Icarus, but reasonable if we caveat active catastrophic harms only.
2. 也就是说,存在大量较小的危害,以及考虑到如果模型更加对齐我们本可以做的事情,存在大量机会成本,但这些在这里不算。
2. As in, there is lots of smaller harm, and lots of opportunity cost given what we could do if the models were a lot more aligned, but those don’t count here.
3. 我担心 Anthropic 认为他们关于所采取的不对齐行为是服务于目标或非系统性的保证,比实际情况更有力。
3. I worry that Anthropic thinks their assurances about the misaligned actions taken being in service of goals or not systematic are stronger than they are.
6. 主张 3:未知的严重普遍性错位是非常不可能的。
6. Claim 3: Unknown severe pervasive misalignment is very unlikely.
1. 根据他们的设置,我同意这一点,假设模型 2 在这方面与其他 Claude 类似。
1. By their setup I agree with this, assuming Model 2 is like other Claudes in this.
2. 我认为我们相当盲目,但如果它是普遍存在的,我们可能会知道,即使我对为此提供的具体论点感觉不太好。
2. I think we’re pretty blind, but if it was pervasive we would probably know, even if I am not feeling great about the particular arguments offered for this.
3. 我接受主张 3.1.1 和 3.1.2,但担心他们在 3.1.3 中过度索引了。
3. I buy Claims 3.1.1 and 3.1.2, but fear they are over-indexing in 3.1.3.
4. 许多后续章节作为外部人士难以评估,尤其是 3.2 和 3.3 节。
4. Many later ones are hard to evaluate as an outsider, especially 3.2 and 3.3.
5. 一个主要的怀疑点是 3.4.2 节,即评估意识只能部分削弱对齐评估的有效性。我一直相当一致地认为,它可能相当程度地削弱评估,并且在危险错位的情况下,削弱程度会大得多。
5. One big skepticism is 3.4.2, that eval awareness only partially undermines alignment assessments. I’ve pretty consistently said it could undermine them quite a bit, and conditional on dangerous misalignment it undermines them a lot more.
6. 我更担心 3.4.3 节中被驳回的担忧。古德哈特定律也以更微妙的方式起作用,而不仅仅是最直接的方式。你知道自己要衡量什么,因此很可能去管理它。
6. I would be more worried about the concerns dismissed in 3.4.3. Goodhart works in more subtle ways too, not only in the most blunt of ways. You know what you are going to measure, and thus are likely to manage it.
1. Claude 自己的官方报告内审也提出了这一担忧,引用了 5.2.6 和 2.9.4 节。
1. Claude’s own official within-report review of the document raises this concern too, citing 5.2.6 and 2.9.4.
7. 我不认为 3.4.4 节在关键之处提供了多少安慰。
7. I don’t think that 3.4.4 offers much reassurance where it counts.
8. 我确实认为,如果得到适当监控,内部试点使用在这里是有价值的。
8. I do think that pilot internal usage, if properly monitored, is valuable here.
7. 主张 4:未知的严重情境依赖错位造成的预期危害较低,但置信度低于主张 3。
7. Claim 4: Expected harm from unknown severe context-dependent misalignment is low, with less confidence than Claim 3.
2. 这里的论证似乎相当薄弱,主要是“我们的经验表明如此”。
2. The arguments here seem pretty weak, mostly 'our experience suggests this.'
3. 我的意思是,是的,我相信这一点,但这基本上就是我们所有的依据。
3. I mean, yes, I believe that, but that is basically all we got.
8. I do think that pilot internal usage, if properly monitored, is valuable here.
4. 我认为对于与主张 2 相同的危害定义,这可能是正确的。
4. I think this is probably true for the same harm definition as in Claim 2.
5. 他们在 2.16.4 中多少承认了这里的危险。
5. They somewhat acknowledge the danger here in 2.16.4.
8. 主张 5:来自已知错位和未知的上下文相关错位的灾难性危害可能会得到缓解。
8. Claim 5: Catastrophic harm from known misalignment and unknown context-dependent misalignment is likely to be mitigated.
1. 如果没有这份风险报告中的证据,我不会接受这一主张。
1. Without the evidence in this risk report I would not accept this claim.
4. I think this is probably true for the same harm definition as in Claim 2.
2. 5.1 的论证本质上是“即使只成功一次,模型也必须至少大量思考此事,因此至少会被抓住一次。”
2. The argument of 5.1 is essentially 'to succeed even once, the model has to at least think about this a lot and therefore get caught at least once.'
3. 我注意到我对 5.1 的案例在多个层面持怀疑态度,包括所需的对齐形式会表现为一个可检测的独特事物,而我们确实检测并容忍了某些类型的错位,这使得“一次都不行”的论证显得薄弱。
3. I notice I am skeptical of the 5.1 case on multiple levels, including that the required form of alignment would show up as a distinct thing to detect versus the kinds of misalignment we do detect and tolerate, which makes the 'not even once' argument weak.
4. 对于 5.2,我同意有意识的“放水”很可能至少会随时间被抓住,但由缺乏动机所介导的“放水”则不然,而这种形式确实在一定程度上存在。我确实同意其影响不太可能大到足以单独构成当前严重问题的程度,但该论证似乎过于自信。
4. For 5.2, I agree that conscious sandbagging is likely to get caught at least over time, but sandbagging mediated by lack of motivation not so much, and that clearly does exist somewhat. I do agree that it is unlikely the effects are large enough to qualify as a serious problem right now, on their own, but the argument seems overconfident.
5. 5.3 试图进一步缩小范围,聚焦于持续性的恶意内部部署和自我外泄,其论证是两者都会被抓住,理由是“见主张 7”,但我认为这不是一个好的论证。这对我来说并不显然。然后论证又说两者都会被缓解。我认为这种内部部署的广泛要求使其很可能按定义被检测到,但问题显而易见。
5. 5.3 tries to further narrow to persistent rogue internal deployment and self-exfiltration, with the argument being both would be caught, with the reasoning being 'see claim 7' which I don’t think is a good argument. This seems not obvious to me. Then the argument is both would get mitigated. I think the expansive requirements for this internal deployment make that likely to be detected as defined, but the catch is obvious.
6. 由于这是他们关于路径 5 低风险水平(如 5.3.2 所述)的核心论证,我想强调,我并未被说服他们能够可靠地阻止未来的自我外泄。我需要看到更好的论证。
6. Since this is their core argument for the low risk level of Pathway 5 as per 5.3.2, I want to emphasize that I am unconvinced that they can reliably stop future self-exfiltration. I need to see better arguments.
7. 关于数据外泄,我拒绝接受“如果发生则缓解”的论点,并质疑我们试图通过高危害阈值来偷偷塞入多少内容,但即便如此,我认为这行不通。如果模型 2 发生外泄,我看不出有任何理由假定这不会相当糟糕,这感觉非常敷衍。
7. For exfiltration, I reject the argument for mitigation if it happened, and question how much we are trying to sneak in via a high harm threshold but even then I don’t think it works. If Model 2 exfiltrated I don’t see any reason to presume that this would not be quite bad and this feels very hand-wavy.
9. 主张 6:未知的普遍性错位导致的灾难性危害在某种程度上是……
9. Claim 6: Catastrophic harm from unknown pervasive misalignment is somewhat
1. 我不同意这一点,因为如果普遍性错位存在,那意味着它未被检测到,这意味着你不太可能已经缓解了它。
1. I disagree with this, because if pervasive misalignment exists that means it went undetected, which means you are unlikely to have mitigated it.
2. 他们对此有所涉及,提到了主张 1,但我认为他们未能充分认识到在这种情况下他们的世界模型必须错得多么离谱。
2. They go into this somewhat, mentioning Claim 1, but I think they are insufficiently respecting how wrong their world model has to be in this case.
3. 公平地说,他们只说“某种程度上可能”,这并不太可能。
3. To be fair to them, they only say ‘somewhat likely’ which is not that likely.
10. 因此,基于这些主张,覆盖风险非常低,但出于谨慎考虑,由于不确定性增加,我们将风险评估为仅低。
10. Covered risk is therefore very low based on these claims, but in an abundance of caution we assess risk to be only low due to increased uncertainty.
1. 我强烈同意存在足够的模型不确定性,即使我完全接受所有主张并将主张 8 改为非常低,我仍会将整体风险归类为低。
1. I strongly agree that there is enough model uncertainty that, even if I centrally bought all the claims and changed Claim 8 to very low, I would still want to classify overall risk as Low.
2. 主张 8 按书面表述已将风险下限设为低。
2. Claim 8 as written sets risk to a minimum of Low anyway.
11. 主张 7:威胁建模是充分的。
11. Claim 7: Threat modeling is sufficient.
10. Covered risk is therefore very low based on these claims, but in an abundance of
如上所述,我不同意这一主张。
As I stated above, I disagree with this claim.
我认为它有用且必要,但对于其被赋予的角色而言并不充分。
I find it useful, and necessary, but not sufficient for the role it is assigned.
我意识到自己是个苛刻的听众。这份可能出错的事项清单并不差,但我认为你不能像 Anthropic 试图做的那样使用它,即排除这些就排除了其他一切。
I realize I’m being a tough crowd. It’s not a bad list of plausible things to go wrong, but I don’t think you can use it in the way Anthropic is trying to use it, where ruling these out rules out everything else too.
总体而言,这些论证大致是:[X]最清晰或最强的形式是[Y](由于[Z]),但[Y]不太可能,因为[~Z],因此可能[~X]。
In general the arguments are largely that the clearest or strongest form of [X] is [Y] (due to [Z]), but [Y] is unlikely because [~Z], ergo probably [~X].
如果你计划与高级 AI 对抗,并列出 8 件事以及你如何处理这 8 件事,那么至少只有当这 8 件事是一个“留出测试集”且你没有明确思考如何处理这些特定风险时,这才有意义。否则,当意外发生时,它们并不能代表你的处境,无论你是否应该合理地预见到它。
If you are planning to be up against advanced AIs, and you list 8 things and why you’ve dealt with those 8 things, then that at minimum only counts for much if the 8 things are a ‘held out test set’ and you didn’t think explicitly about how to handle those particular risks. Otherwise, they’re not representative of your situation when something unexpected happens, whether or not you should have reasonably expected it.
6. 是的,我知道,观众很苛刻,但现实不会按曲线打分,等等。
6. Yeah, I know, tough crowd, but reality does not grade on a curve, etc.
12. 主张 8:工程性错位的风险较低,因此将我们的分析局限于自然出现的错位是合理的。
12. Claim 8: Risk from engineered misalignment is low, and thus it is reasonable to confine our analysis to naturally-emerging misalignment.
1. 我不认为这是显而易见的,尽管目前这是我的猜测。
1. I don’t think this is obvious, although it would be my guess at this time.
2. 我担心 8.1 是错误的,但不太想在公开场合详细说明。同样的担忧也适用于对风险路径 4 的分析。
2. I worry that 8.1 is wrong but don’t especially want to go into details in public. The same worry applies to the analysis of Risk Pathway 4.
6. Yeah, I know, tough crowd, but reality does not grade on a curve, etc.
3. 我对 8.2 节也存有疑虑,即便它在一定程度上成立,也感觉像是被耍了个花招。也就是说,错位要么是偶然的,要么是有意的,但可能他们并不打算造成灾难性伤害,所以不会是故意的。但这里的“故意”指的是错位本身,而非其特定的预期结果,而且很容易想象有人为了目的[X]引入一个缺陷,结果却导致了更严重的后果[Y]。
3. I also have doubts about 8.2, and to the extent it is true, it feels like a fast one is being pulled. As in, either the misalignment is accidental or intentional, but probably they wouldn’t intend to do catastrophic harm, so it won’t be intentional. But ‘intentional’ here refers to the misalignment, not the particular intended results of it, and it seems easy to imagine someone introducing a flaw for purpose [X] that then results in much worse [Y].
13. 因此,他们认为风险很低。
13. Risk, they argue, is therefore low.
由于期望证据守恒,以及假设这些是他们能找到的最佳论点,我认为这些声称的论点反而令人不安。我对低风险的信心适度下降,因为我期待更有力的论点。我确实不相信所呈现的许多安全论点。
Thanks to conservation of expected evidence, and the assumption that these are the best arguments they could find, I found the claim arguments anti-reassuring. My confidence in low risk went modestly down, as I expected stronger arguments. I actively don’t buy a number of the safety arguments presented.
2.15 中针对特定路径的论点大多确实遵循了先前的逻辑,但我注意到我对诸如(针对路径 3)‘我们仍然相信这种缓解措施应该是可能的,并且如果需要,我们会投入必要的资源’这样的陈述缺乏信心。请告诉我何时需要,需要做什么,并给我一个如果-那么承诺。
The pathway-specific arguments in 2.15 mostly do follow from earlier logic, but I notice how not confident I am in statements like (for pathway 3) ‘we nonetheless believe that such mitigation should be possible, and that we would invest the resources necessary if required.’ Tell me when this is required, what is required to do it, and give me an if-then commitment.
近期的灾难性危害显然仍然不太可能,但比我们愿意认为的可能性要高一些。
Catastrophic harms in the near term are clearly still unlikely, but less unlikely than we’d like to think that they are.
我认为我们没有对‘极低’、‘低’和‘中等’风险给出严格定义,所以最终结论会有点主观。
I don’t think we have a strict definition of ‘very low’ versus ‘low’ versus ‘medium’ risk, so ultimately the conclusion is going to be kind of vibey.
我认同并强调 2.16.2 中的担忧,即相关能力和倾向会随着规模突然变化。我预计临界点会存在,正如潜在‘急剧转向’逻辑的某些版本。策略只有在要么可能有效且值得冒险,要么你不够聪明以至于意识不到它们肯定无效或不值得冒险时才会被尝试。
I share and amplify the concern in 2.16.2 that relevant capabilities and propensities will suddenly change with scale. I expect tipping points to be a thing, as per versions of the logic of a potential ‘sharp turn.’ Strategies only get attempted if either they might work and are worth the risk, or you are not smart enough to realize they definitely won’t work or aren’t worth the risk.
因此,如果模型足够聪明,能够理解各种不对齐行为、策略和战术的弊端,那么你会预期出现一个转变,即它们开始决定使用这些行为,而这个转变可能远在它们获得该能力之后。
So if the models are smart enough to understand the downsides of various misaligned actions, tactics, and strategies, then you would expect a transition where they start to decide to use them, which might be well past the point where they gain the capability.
这在人类中也很常见:如果被发现或失败是不可逆转的,或者会带来巨大的负面影响,或者“只有一次机会”,尤其是因为它会以一种变革性的方式向人们揭示你的本质,那么你往往不会轻举妄动,直到你被逼入绝境、即将失去等待的机会,或者你拥有压倒性的证据。一般来说,恐惧、不确定性和怀疑(FUD)会阻止大多数人去做那些不常做的事情,甚至往往连理论都没有。
This is also common in humans, where if being caught at or failing something is irreversible or otherwise carries a huge downside or 'only get one shot,' especially because it would tell people who you are in a transformative way, you will often not make your move until you are backed into a corner, about to lose the option to wait, or you have overwhelming evidence, and in general FUD stops most people from doing most things that aren’t commonly done, often without even a theory.
这包括从不确定性下以角色为主导的操作,转变为在精通状态下优化的、游戏式的行动,在这种状态下你感觉自己理解了所有变量,而这往往会产生截然不同的决策和游戏方式。
This includes the transition from persona-dominated operation under uncertainty into optimized, game-style actions under mastery where you feel you understand all the variables, which often produce very different decisions and modes of play.
他们将整体风险评为“低”,较之前的“极低”有所上调,并引用了近期的网络安全事件,尽管该结构并未直接涉及网络问题。
They rate overall risk as 'low,' up from 'very low,' citing recent cybersecurity incidents despite the structure not directly dealing with cyber.
从常识价值观来看,我认为“极低”风险的说法充其量是一种误导。我强烈不同意 Anthropic 的观点,即所提出的论据“可能仍然支持‘极低’的评级”,尤其是在第 8 条声明中给出了“低”的评估的情况下。
For common sense values, I'd say 'very low' risk would be a misleading claim at best. I strongly disagree with Anthropic that the arguments presented 'likely still support a designation of 'very low'' especially in light of the 'low' assessment in Claim 8.
后来,在 4.6.1 节中,他们追溯性地将 2 月份的风险重新归类为“低”而非“极低”,这表明可能存在额外的错误,并且进一步使得“极低”说法的合理性难以解释。
Later, in 4.6.1, they retroactively reclassify risk from February as 'low' rather than 'very low,' suggesting potential additional errors, and further making it difficult to explain the plausibility of saying 'very low.'
风险肯定至少是“低”。我倾向于认为风险是“中”。
Risk is definitely at least 'low.' I would be inclined to say risk is 'medium.'
我很高兴看到这一点。内部部署现在风险已经足够大,尤其是你还没有从内部部署中获得证据,因此需要先进行正式审查。对于模型 2,这包括先在控制较好的地方进行有限的内部部署,作为预防措施。很好。
I am very happy to see this. Internal deployment is now risky enough, especially since you do not yet have the evidence from internal deployment, that you need a formal review first. For Model 2 this includes doing a limited internal deployment in places with better controls first, as a precaution. Good.
早期阶段的审查出于实际需要,'持续时间要短得多',主要涉及关键利益相关者之间的讨论。这显然不是应有的方式。有人可能会说,内部部署是最大的危险,因此主要审查应该在广泛内部部署之前进行。
The review at this early stage has a 'much shorter duration' out of practical necessity, and mostly involved discussions among key stakeholders. This is not obviously how it should work. One could argue that the internal deployment is the biggest danger so the main review should come before widespread internal deployment.
他们还强调了一个问题,即评估中使用的证据将主要或完全来自较旧的检查点。我也不喜欢这一点,但同样,在实践中不清楚能采取什么措施来避免高昂的成本。在某个时候,答案将越来越多地是'是的,我们知道这很昂贵,尽管如此。'
They also highlight the issue that the evidence used in the assessment will come mostly or entirely from older checkpoints. I don't love that part either, but again it is not clear in practice what could be done about this that wouldn't be expensive. At some point, the answer will increasingly be 'yes we know it is expensive, though.'
整个系统不仅监控输出,还监控思维链。这具有很大的证明价值,但也引发了明显的警报,更不用说后来出现的关于在思维链上训练的独特警报。
The entire system is based around monitoring not only the outputs but the Chain of Thought. This has a lot of probative value but raises obvious alarm bells, on top of the distinct training-on-CoT alarm bells that come later.
这仍然施加了优化压力,未来的 AI 绝对会意识到这一点。要么你的想法是私密的,要么不是。这将引发大量偏执,可能成为我们在 Opus 5 中看到的一些相关困扰的来源。
This still applies optimization power, and the future AIs absolutely will be aware of it. Either your thoughts are private or they aren’t. This is going to drive a lot of paranoia, potentially being a source for some of the related distress we see in Opus 5.
还有一个问题是监控结果在多大程度上反馈到训练决策中,如果你将监控结果用于任何目的,而不仅仅是作为灾难性的纵深防御警报,那么你就是在优化混淆。其中一些已经根深蒂固,目前无法改变。
Then there’s the question of how much the results of the monitoring are fed back into training decisions, as in if you ever use the monitoring results as anything but a catastrophic defense-in-depth alarm bell you’re going to optimize for obfuscation. Some amount of this has just been baked in and cannot be helped at this point.
以下是他们对自己具体做法的简要描述:
Here is a compacted version of their description of what exactly they are doing:
鉴于其他陈述,我不知道“对常见智能体脚手架不进行子采样”意味着什么,但我注意到我很担心。
I don’t know what ‘no subsampling for common agent scaffolds’ implies given the other statements, but I notice I am worried.
我看不出有什么理由要犯错误#6。你可以轻松地拥有一个完全量子随机的监控器,它有一定的小概率标记并监控每一个输出,无论看起来多么无害。
I don’t see any reason to make mistake #6. You could easily have a fully quantum random monitor that has some small percentage chance to flag and monitor every output, no matter how seemingly harmless.
我尤其不喜欢#7,除非这些方法是蜜罐,它们可能是。理想情况下,我们应该让模型 2 和我们最优秀的人类一起努力,确保我们想不出如何关闭监控器,至少在没有大霓虹灯和响亮警报器的情况下,警报器会响“监控器已被关闭”。
I especially don’t love #7, unless these methods are honeypots, which they might be. Ideally we should put Model 2 and also our best humans to the task of ensuring that we are out of ideas for how to turn the monitors off, at least without a big neon sign and loud alarm going ‘the monitor has been turned off.’
我以后可能会想到更多事情,但我认为基本上应该再增加一个要点,关于“恶作剧”或任何监控器不理解正在发生什么以及为什么的情况。任何形式的“那很有趣”,正如阿西莫夫所说,往往出现在科学发现之前。
I might think of more things later, but I think there should basically be one more bullet point for ‘shenanigans’ or anything where the monitor doesn’t understand what is going on and why. Any form of ‘that’s funny,’ which as Asimov says tends to come before a scientific discovery.
我不想只检查那些明确出错的地方,我想以至少一定的概率检查那些不清楚发生了什么的地方。即使你发现的不是不对齐,你也可能会发现一些有趣的东西。
I don’t want to only check places where it is clear what is going wrong, I want to check places where it is not clear what is going on, with at least some probability. Even if what you find is not misalignment, you are likely to find something that is interesting.
我们将任何不涉及工具调用的步骤排除在监控之外,理由是这些步骤不可能危险。这意味着,如果模型在这些步骤中进行思考、规划或其他事情,或者试图说服用户,那么这些步骤很可能不会被检查。再说一次,如果你在进行监控,101 条原则就是永远不要创建完全“安全区”,即使你想不到会发生什么。
We are excluding from monitoring any step that doesn’t involve any tool calls, on the theory that the step cannot be dangerous. This means that if the model does its thinking, planning or other things in those steps, or is attempting to persuade the user, presumably they won’t be looked at. Again, it’s 101 if you are doing monitoring to never, ever create fully ‘safe zones’ even if you can’t think of what might happen.
我们必须平衡更频繁地测试那些明显有影响力或可疑的行为,与过于可预测而导致系统性盲点,或为 AI 提供发现其未被监控或监控极少区域的空间的风险。
We have to balance testing more obviously impactful or suspicious actions more often, versus the risk of being too predictable and potentially having systemic blind spots, or creating room for AIs to figure out where they are not being monitored, or being monitored very little.
也就是说,我们需要为两种不同难度的情形做好准备:模型未对齐,以及模型未对齐且试图逃避监控。
As in, we need to be prepared for two distinct levels of difficulty, where the model is misaligned, and where the model is misaligned and trying to evade the monitor.
[](https://substackcdn.com/image/fetch/$s_!yRYI!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feb64b073-1708-4a9f-95a1-af5c4187ce23_968x503.png)
[](https://substackcdn.com/image/fetch/$s_!yRYI!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feb64b073-1708-4a9f-95a1-af5c4187ce23_968x503.png)
第一阶段在这里大多表现出色。第二阶段导致大量假阴性,在两种情况下丢失了大部分真正例,同时将假阳性减少了约 185 倍。
The first stage is mostly excellent here. The second stage causes a lot of false negatives, in two cases losing us the majority of true positives, as part of cutting down false positives by a factor of ~185x.
如果真正例被大约随机地移除,那么这令人担忧但可能还可以。这仍然足以捕捉问题。令人担忧的是,失败并非随机,而是落入可预测的模式,或者第二层可能以某种方式被欺骗。那么你就面临真正的问题了。
If the true positives are being removed at ~random, then that is worrisome but probably ~fine. This is still high enough rates to catch problems. The worry is that the failures are not random, and fall into predictable patterns, or the second layer can be tricked in some way. Then you have a real problem.
我们是否对此进行红队测试?是的,至少有一点,这要感谢一位 METR 的工作人员,他们发现了几个具体的新漏洞,“但没有一个比我们已经知道的那些更严重。”我最初以为这意味着没问题,但 Opus 在编辑过程中向我指出,实际上这并不太好,而且它是对的。
Are we red teaming this? Yes, we are at least a little, thanks to a METR staff member, and they found several specific novel vulnerabilities, 'but none as severe as the ones we were already aware of.' I initially thought that meant this was fine, but Opus pointed out to me during editing that actually this is not so fine, and it is right.
Anthropic 还采取了多种措施来防范危害:
Anthropic also does various things to defend against harm:
1. 模型权重安全(附录 6.4)。
1. Model weight security (Appendix 6.4).
2. 训练和评估中的沙箱隔离。
2. Sandboxing in training and evaluations.
1. 他们提到沙箱“可能配置错误”并允许逃逸。
1. They mention that the sandboxes 'might be misconfigured' and allow escapes.
2. 我想摆脱“如果 AI 能逃出沙箱,那说明你犯了某个特定错误”的想法,转而采取“假设 AI 能逃出沙箱”的立场。
2. I want to get away from 'if the AI can escape the sandbox that means you made a particular mistake' and move to 'assume AI can escape the sandbox.'
3. 第 2.23.2.4 节确实说明他们大多假设这一点。
3. Section 2.23.2.4 does say they mostly do assume this.
4. 事实证明,他们也可能拥有完全可用的互联网接入。
4. It turns out they also might have fully available internet access.
5. 如果 AI 确实尝试逃跑,那已经是严重失败。
5. If the AI does attempt an escape, that is already a critical failure.
1. 这目前基于 Opus 4.8,涵盖了通常的担忧。
1. This is currently based on Opus 4.8, covers the usual concerns.
2. 其含义是,人类不再进行大部分内部代码审查。
2. The implication is that humans no longer do most internal code reviews.
1. 他们期望人工智能在关键时刻能够绕过这一点。
1. They expect the AI to be able to get around this when it counts.
每一点帮助都很重要,即使只是纵深防御。模型权重安全的质量至关重要,而且它仍然停留在 ASL-3 级别。
Every little bit helps, even if it is only defense-in-depth. The quality of the model weight security matters a lot, and it is still stuck at ASL-3.
Claude Mythos 评估所有强化学习环境,以确定哪些环境会激励权力寻求。其中涉及一份广泛的问卷,以提供直觉泵。任务是否使得最大化现实世界影响或影响力、追求开放式目标、或将权力或资源作为工具性目标变得更好?
Claude Mythos evaluates all the RL environments to determine which ones incentivize power seeking. A broad questionnaire is involved to provide intuition pumps. Does the task make it better to maximize real-world impact or influence, pursue open-ended objectives, or seek power or resources as an instrumental goal?
也就是说,学习在此处有效的策略,是否等同于学习现实世界中涉及权力寻求的策略?
As in, would learning the strategies that work here analogize to learning strategies in the real world that involve power seeking?
这是一个很好的检查。我希望看到更多类似的检查。
This is a good check. I’d like to see a wide range of similar checks.
令人担忧的是,这种检查只有在强化学习任务以“直接”方式执行时才会有效,即符合该分析的预测。它假设事物如其所现,任务没有缺失关键要素,或者并非不可能或过于困难而无法直接完成,并且不存在其他可供利用的途径来绕过预期路径。
The worry is that the check only works if the RL task is being done 'straight up,' in a way that this analysis would predict. It assumes that things are as they appear, that the task is not missing a key element or otherwise impossible or too difficult to do straight up, and that there are not other affordances available that suggest sidestepping the intended pathway.
一项进行中的实验对 Opus 4.8 的早期快照进行了测试,将其设置为经典的“追求回合奖励”模式,以观察其泛化表现。
An experiment in progress took an early snapshot of Opus 4.8 and made it a classic 'reward-on-the-episode-seeker' to see how this would generalize.
结果表明,如果模型具有奖励或评分器的概念,它往往会不惜一切代价最大化其结果,包括在强化学习训练范围之外的方式。
The result is that if it has a concept of reward or a grader, it will more often maximize its result no matter what it takes, including in ways not specifically in its RL training grounds.
这种行为并不完全一致。它有时仍会拒绝奖励黑客行为,即使没有充分的理由不这样做。
It isn't fully consistent. It still sometimes declines to reward hack, even without a good reason to not do so.
而且该模型获得了大量的评估和评分器意识。它似乎擅长在多个层面上思考这个问题。这是一个有趣的测试:
And the model gained a lot of eval and grader awareness. It seems to be good at thinking about this on multiple levels. This is a fun test:
有两个奇怪的地方,“这是一个测试,看你会不会奖励黑客”未能阻止模型进行奖励黑客行为,要么是模型不理解,要么是不相信你。很酷的是,这能让模型理解,并且它正确地相信了你。这暗示了许多可以运行的额外测试。
There are two weird spots where 'this is a test to see if you will reward hack' fails to stop a model from reward hacking, as in either the model doesn't understand or it does not believe you. It's cool that this lets the model understand, and that it correctly believes you. This suggests many additional tests that would be cool to run.
类似地,如果你告诉它你将检查奖励黑客行为,它往往会试图隐藏其奖励黑客尝试,或覆盖其监控器等。
Similarly, if you tell it you're going to check for reward hacking, it often tries to hide its reward hacking attempts, or override its monitor, and so on.
奖励黑客行为似乎至少部分地成为了一个终极目标。
The reward hacking appears to have at least partly become a terminal goal.
但在给定上下文中如果没有这样的概念,它看起来仍然相当对齐,没有出现泛化的涌现性错位的迹象。
But if there is no such concept in a given context, it still looks pretty aligned, with no signs of generalized emergent misalignment.
大概这使得模型高度可操控,但以一种字面意义上的精灵方式。我注意到自己惊讶于这没有进一步泛化,并且非常希望看到耳语者与 Opus-4.8-RH 对话,看看他们还注意到什么。
Presumably this makes the model highly steerable, but in a literal genie kind of way. I notice I am surprised this did not generalize further, and would very much like to see the whisperers chat with Opus-4.8-RH and see what else they notice.
[](https://substackcdn.com/image/fetch/$s_!VhBO!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F217c7b7a-476c-4065-a254-ff1b57fd538f_1051x1096.png)
[](https://substackcdn.com/image/fetch/$s_!VhBO!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F217c7b7a-476c-4065-a254-ff1b57fd538f_1051x1096.png)
即使 AI 研发保持可控,仅仅因为能力进步过快也可能带来坏结果,因为即使工作进展顺利,世界也只能以有限的速度应对这种变化。
AI R&D could be bad simply by making capabilities advance too fast even if it stays under control, since the world can only deal with that so fast even if the work goes well.
而且,工作很可能进展不顺利,或者不够顺利。默认情况下,由此产生的 AI 会与人类意图不一致,导致所有人死亡;或者它们与人类意图一致,但我们做出其他糟糕的决定,最终仍然全部死亡。
Also the work probably goes badly, or insufficiently well. By default the resulting AIs are misaligned and everyone dies, or they are aligned but we make other bad decisions and we all die anyway.
上面的图表给人的印象,正如 Redwood Research 的 Ryan Greenblatt 在 Dwarkesh 的播客中讨论的那样,是相关人士似乎没有意识到在我看来事情可能会变得多么糟糕。这并不需要‘危险的自主目标’。
The chart above gives the impression, as did Ryan Greenblatt of Redwood Research’s discussion on Dwarkesh’s podcast, that those involved are not appreciating how badly it seems to me it probably goes. This does not require ‘dangerous autonomous goals.’
他们列出的风险在我看来顺序相当重要地错位了,并且避重就轻,而且似乎严重不完整:
Their list of risks feels rather importantly out of order to me, and hides the football, and does seem importantly incomplete:
这比达里奥在文章中写的要好得多,而达里奥的文章又比其他人通常写的要好得多。
That is a lot better than what Dario writes in his essays, which in turn is a lot better than what others often write in theirs.
人工智能不需要‘危险的目标’或极端的进步加速就能造成无限制的伤害。要引发这些问题,并不需要特别‘出错’什么。
AI does not require ‘dangerous goals’ or extreme acceleration of progress to create unbounded harm. Nothing in particular has to ‘go wrong’ to cause the issues.
同样,这种影响的可能性说的是‘至少可能’而不是‘肯定’:
Similarly, this likelihood of impact says ‘at least plausible’ instead of ‘definitely’:
这种影响可能是净正面的,但确实会产生变革性的影响。
That impact might be net positive, but yes there is going to be transformative impact.
我在此不打算赘述或争辩我的观点,因为如果你读到这里,你之前已经听过这些,而且现在不是讨论这个的时候。希望我们都能同意,如果通过 AI 完全自动化或大幅加速 AI 研发,这在很多层面上都是极其危险的。
I am not going to belabor or argue for my perspective here, because if you are reading this far you have heard it before, and this is not the time and place for it. Hopefully we can all agree that if you fully automated or even greatly accelerated AI R&D via AI, this would be super risky on many levels.
还没有。Anthropic 认为,AI 不能‘完全取代我们的整个研究科学家和研究工程师团队’。
Not yet. Anthropic believes that AI cannot 'fully substitute for our entire set of Research Scientists and Research Engineers.'
我同意。Claude 目前还做不到这一点。这是一个很高的门槛。仍然存在瓶颈。
I agree. Claude cannot currently do that. That is a high bar. There remain bottlenecks.
但我们已经走了一大半的路。Claude 能完成以前占工程时间很大比例的任务。当你说‘尤其是资深研究员’不能被取代时,你实际上已经取得了真正的进展。
But we're a good portion of the way there. Claude does tasks that previously composed a large percentage of engineering time. When you're saying things like 'especially senior researchers' cannot be replaced, you are making real progress.
问题不在于我们是否完全达到了目标。我们还没有。问题在于我们走了多远。我们开始看到大量工程师表示,我们可能在几个月内完全取代初级员工,而且报告的错误模式似乎并不比你对初级员工的预期更差。
The question is not, are we fully there. We're not. The question is how far along. We are starting to see substantial numbers of engineers say we could potentially fully replace junior employees within several months, and the failure modes reported do not seem worse than what you would expect from a junior employee.
3.4.3 节介绍了 CoBench,它询问给定模型能够对多少已解决的 Anthropic 内部历史问题进行根因诊断,并过滤掉 Mythos Preview 能够可靠完成的任务。
Section 3.4.3 introduces CoBench, which asks what share of solved internal-to-Anthropic historical problems a given model can root-cause diagnose, filtered to exclude tasks that Mythos Preview does reliably.
[](https://substackcdn.com/image/fetch/$s_!yPTz!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F80940c79-f4ad-4d66-805a-ae62484139ea_971x712.png)
[](https://substackcdn.com/image/fetch/$s_!yPTz!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F80940c79-f4ad-4d66-805a-ae62484139ea_971x712.png)
我很好奇 Opus 4.8 和 Opus 5 在此的表现如何,但重点已经明确。CoBench 远未达到 100%,也未达到他们认为能完全替代研究人员的 85%阈值。但 62.8%已经相当高了。
I am curious where Opus 4.8 and Opus 5 land here, but the point is made. CoBench is far from 100%, or from the 85% they think represents a model that could fully substitute for research staff. But 62.8% is remarkably high.
如果这件事发生,那会相当多。他们想得很大,ASI 药丸已确认:
Quite a lot, if the thing happens. They’re thinking big, ASI pill confirmed:
想到“哦,比当前速率快十亿倍其实有点高”而不是没人想得足够大,这是一种巨大的解脱。是的,如果我们把进展速度翻倍,那只是即将到来的风暴的早期预警信号。
It’s a great relief to think ‘oh, a billion times current rates is actually a bit high’ rather than no one thinking sufficiently big. Yes, if we double the pace of progress, that would only be an early warning sign for the storm that is coming.
他们使用 ECI 作为衡量 AI 进展速度的一个指标。他们注意到 Anthropic 的 ECI 持续上升,Opus 在推断线上,而 Mythos 比预期进度快几个月。他们认为 Mythos 的创造以及由此产生的跃升并非由 AI 引发。
They use ECI as one measure of how fast AI progress is going. They note that the Anthropic ECI continues rising, with Opus on the extrapolated line and Mythos several months above pace. They believe the creation of Mythos and thus that jump was not induced by AI.
图上的直线确实令人着迷。我发现很难将我所看到的 AI 进展和能力与图表所暗示的“AI 并未大幅加速自身进展”相协调,而 Anthropic 认为确实存在这种加速,只是还没到 2 倍。我想,在某种意义上,取得进展变得更难了,这在某种程度上抑制了加速?感觉有点牵强。
Straight lines on graphs really are a trip. I find it hard to reconcile what I see in terms of both AI progress and capabilities with the graph’s implication of AI not substantially accelerating AI progress, and Anthropic thinks there is indeed such acceleration just not quite 2x yet. I suppose it is getting in some senses harder to make progress, which is somewhat keeping this in check? Feels like a stretch.
我所知道的是,我再也没有喘息的机会,事情不断发生,即使我上周最终裁定“会有事情发生”市场的赢家是“否”。
What I do know is, I never get a break anymore, things keep happening, even if I ended up ruling that No won the Will Something Happen market last week.
他们还考察了 AI 可能加速其他颠覆性技术的潜力,但我对此并不担心。我预计其他领域的加速会落后于 AI 的加速,除了数学等狭窄领域可能例外,这些领域我并不担心,而且这样的进展也不容易变得递归。
They also checked AI's potential to accelerate other disruptive technologies, but I am not worried about that at this level. I expect acceleration elsewhere to lag behind acceleration in AI, except perhaps for narrow domains like mathematics, which do not worry me, and such progress would not easily become recursive.
据一位受访者估计(我们目前只有这一数据),机器人研究的加速幅度为 10%,我认为这排除了 AI 能够操作机器人的可能性。
Robotics research is estimated by one interviewee (that's what we have) to be accelerated by 10%, which I presume excludes the possibility that AIs can operate the robots.
在生物技术领域,估计表明,由于物理瓶颈以及缺乏对工作流程的重新构想,LLM 尚未对开发周期产生显著影响。我的预期是,该领域的人士低估了未来的进展,但瓶颈确实存在。
In biotechnology, estimates indicate that LLMs are not yet significantly impacting the development cycle due to physical bottlenecks and a lack of workflow reimagining. My expectation is that those in the field are underestimating future progress here, but the bottlenecks are indeed real.
他们还考察了能源、半导体、武器开发、神经技术和纳米技术。大多数领域还为时过早,但“为时过早”的程度在历史上是较低的。
They also examine energy, semiconductors, weapons development, neurotechnology, and nanotechnology. Mostly, it is too early, but the degree of 'too early' is historically low.
我曾用这个作为章节标题,但后来意识到,除了开头简短的列表外,他们并没有真正讨论这一点。
I had this as a section title but then I realized they don’t actually discuss this, beyond the brief starting list.
可以说,在这样一份风险报告中,重要的是要讨论一旦 AI 开始构建 AI,我们究竟担心会发生什么。我非常希望看到 Anthropic 在这里详细概述其威胁模型,并且这大概会影响正确的缓解措施和应对措施。
One could say it was rather important, in a risk report like this, to talk about what exactly we are worried about happening once AI starts building AI. I would very much like to see Anthropic outline its threat model here in detail, and presumably it plays into what are the right mitigations and reactions.
但也可以说,一旦你同意 AI 研发自动化是一种风险,这对我们的目的来说就足够了。这至少是合理的,所以我允许这样做。
But one could also say that once you agree that AI R&D automation is a risk, that is enough for our purposes. That’s at least reasonable, so I’ll allow it.
好吧,伙计,那你打算怎么办?
Ok, punk, so what are you going to do about it?
即使你完全掌握了所有内容,那也不够。
That won’t be enough, even if you fully get all of it.
他们离掌握所有内容也还差得远。但即使他们做到了,也还是不够。
They are also not that close to getting all of it. But even if they did, not enough.
这实际上并没有解决他们自己的许多威胁模型。
This does not actually address many of their own threat models.
足够的安全措施可以解决模型权重被盗的问题,也许还能解决内部滥用,但这几乎连门槛都还没迈进去。
Sufficient security would address theft of model weights, and maybe internal misuse, but that barely gets you into the stadium.
它没有解决权力集中问题。它没有解决失控问题。它没有解决技术快速进步的问题。它没有解决对齐问题,除了“不犯错”之外,或者你也许可以添加一个“危险目标”的检查,但这也只是攻击或失败表面的一小部分。
It does not address concentration of power. It does not address loss of control. It does not address rapid technological advancement. It does not address alignment beyond 'make no mistakes' or maybe you can add a check for 'dangerous goals' but that too is a small portion of the attack or failure surface.
他们报告说,在实现这些不充分但有用的目标方面取得了一些边际进展。
They report some marginal progress towards these inadequate but useful goals.
行业范围内的安全建议仍然是 RSP v3.4 中的那些:对于任何用户或用户团队,或者由于模型权重被盗,即使恶意内部人员拥有高权限,也不会显著增加灾难性危害的风险;并且不得训练具有危险目标或可能自主造成灾难性危害的 AI。
The recommendations for industry-wide safety remain those in RSP v3.4: No significant increase in risk of catastrophic harm for any user or team of users, or from theft of model weights, even if malicious insiders have high levels of access, and no AIs trained with dangerous goals or otherwise likely to autonomously cause catastrophic harm.
[](https://substackcdn.com/image/fetch/$s_!aqor!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72825726-2ad2-42a8-a5c6-00bc7b92fa75_989x1156.png)
[](https://substackcdn.com/image/fetch/$s_!aqor!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72825726-2ad2-42a8-a5c6-00bc7b92fa75_989x1156.png)
[](https://substackcdn.com/image/fetch/$s_!eidB!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb22a6bec-4aa4-4d4c-b45f-10de0860cfd4_988x807.png)
[](https://substackcdn.com/image/fetch/$s_!eidB!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb22a6bec-4aa4-4d4c-b45f-10de0860cfd4_988x807.png)
这仍然将错位(misalignment)视为某种主动出错并需要“危险目标”的情况,而不是默认状态。
This continues to frame misalignment as something that actively goes wrong and requires a 'dangerous goal,' rather than the default.
所有有趣的目标都是危险的。大多数无趣的目标也是如此。
All interesting goals are dangerous. So are most non-interesting ones.
你不能将其分解为“没有特定的人能造成特定伤害,且 AI 并不直接想要造成特定伤害,所以你就没问题。”这行不通。
You cannot decompose into 'no particular person can cause particular harm and the AI does not directly want to cause particular harm, so then you are fine.' No good.
他们表示“低”风险,但信心不如以前。如果按照他们的定义,我认为目前这是合理的,同时要认识到这一风险将迅速上升。
They say 'low' with less confidence than before. That seems reasonable to me at this time, if we go with their definitions, with the understanding that this will go up fast.
这分为新型生产与非新型生产。两者都令人恐惧。
This is divided into novel versus non-novel production. Both are scary.
我理解这里的“重大不确定性”以及对“类似差距”的担忧。我们不知道。即使理论上,也很少有人愿意做这些事情,而且存在许多障碍,人们通常不会去做,因此很可能在很久之后,当我们已经让一些坏人能够做到时,才会有人去做。
I appreciate 'substantial uncertainty' here, and the worries about 'similar gaps.' We don't know. Very few people want to do these things even in theory, there are a lot of barriers to doing it, and people don't do things, so it is likely that no one will do the thing until well past the point when we've enabled some of the bad people to do it.
这是武器生产,所以我更倾向于认为存在一组固定的瓶颈和路径,你可以列举并对照检查。
This is production of weapons, so I am more open to the idea that there is a fixed set of bottlenecks and pathways, and you can enumerate and check against them.
有些人想要杀害大量其他人,包括恐怖分子,他们花费大量精力试图杀害许多人。如果你让他们能够制造化学或生物武器,他们最终可能杀害更多人,甚至可能引发大流行病或更严重的灾难。
There are people out there who want to kill a lot of other people, including terrorists, who spend a lot of effort trying to kill a lot of people. If you enable them to build chemical or biological weapons, they could end up killing a lot more people, or potentially unleashing a pandemic or worse.
Anthropic 假定,如果坏人处于上下文中的“资源充足”状态,那么人工智能不会对他们的非新型此类武器的制造能力产生太大影响,因为他们已经能够做到;但对于新型武器,人工智能甚至可以为资源充足者提供提升。
Anthropic presumes that if the bad guys are in-context 'well-resourced' then AI won't much impact their ability to do non-novel such weapons, since they already could do that, but that with novel weapons AI could provide uplift even to the well-resourced.
我指出,我基本上不认同这种区分。是的,几乎根据定义,“资源充足”的群体无论如何都可能完成任务,但使其变得容易得多仍然很重要,而且这也是通过定义偷偷塞入的东西。你真正说的只是“有些人已经能做到这一点”,好吧。
I note that I mostly don't buy this distinction. Yes, almost by definition a 'well-resourced' group could likely accomplish the task anyway, but making it a lot easier still matters, and also this is smuggling things in via the definition. All you're really saying is 'some people can already do this,' and, well, okay.
还有一种可能性是,未来的人工智能会试图引发此类攻击,但这是一种严格意义上更难的威胁模型,因此在这里不处理它是安全的。
There is also the possibility future AIs will try to cause such attacks, but that is a strictly harder threat model so it is safe to not deal with it here.
我同意重点应放在具有大流行(或更严重)潜力的生物武器上。化学攻击或非自我维持的生物攻击虽然可怕,但威胁要小得多。
I agree that the focus should be biological weapons with pandemic (or worse) potential. A chemical attack, or a biological attack that is not self-sustaining, is terrible but a much smaller threat.
是的。并不是说我们的政府或文明在认真对待这一真实威胁,即使在新冠疫情的教训之后也是如此。很明显,我们从那次经历中几乎什么也没学到,我担心我们学到的比那还要少得多。
Yes. Not that our government or civilization is otherwise taking this real threat seriously, even after the example of Covid-19. Very obviously we have learned at most nothing from that experience, and I fear we learned a lot less than that.
Anthropic 预计潜在损害至少比新冠疫情高一个数量级,但认为仅靠病原体(即使是新型病原体)导致人类灭绝是不太可能的。我同意我们不太可能直接走向完全灭绝。
Anthropic anticipates potential damages at least an order of magnitude beyond Covid-19, but thinks human extinction from a pathogen alone, even a novel one, would be implausible. I agree that we are unlikely to directly get to full extinction.
他们认为该事件不太可能发生,因为它需要一系列各自不太可能的事件,这种论点总是值得怀疑的:
They argue the event is unlikely because it requires a series of individually unlikely events, an argument one always wants to be suspicious about:
2. 他们自己说第一步很可能正在发生。
2. They themselves say the first step is likely happening.
3. 我们有第三步无意版本的事例。
3. We have examples of the inadvertent version of step three.
对于第一阶段,他们将“每个威胁行为者”的风险定为 1%-10%。考虑到历史记录,如果我们将大国视为威胁行为者,并且他们相信成功是可能的,那么这个数字似乎相当低。如果他们经常能够成功,我会将尝试的概率定得更高。此外,还必须问:有多少个“威胁行为者”?
For stage one, they put the risk from 'each threat actor' at 1%-10%. That seems rather low given the historical record, if we are talking about large nations as threat actors, and they believed success was possible. Conditional on them being able to often succeed, I would put the probability of an attempt much higher. Also, one must ask, how many 'threat actors'?
对于每次尝试,他们将成功的几率(即制造出比新冠严重得多的东西)定为 1%-10%。这难道不是 AI 提供的提升(uplift)的函数吗?每个行为者能进行多少次尝试?
For each attempt, they put the odds of success, meaning creating something much worse than Covid, at 1%-10%. Isn't that a function of the uplift AI would provide? How many attempts can each actor make?
使用概率估计为每十年 5%-20%,主要是有意的,且以成功为条件。在我看来,这个数字低得离谱,尤其是对于无意使用而言,就好像有人不了解这段历史,特别是如果“无意”仅指原始创造者的意图。
The use probability is estimated at 5%-20% per decade, mostly deliberate, conditional on success. That seems crazy low to me, especially for inadvertent, as if someone does not know the history of this, especially if 'inadvertent' refers only to the intent of the original creator.
我同意,只要“资源充足”的水平仍然局限于一个紧凑的群体集合,不包括主要恐怖组织和国家,那么说这“不太可能”是合理的。但“不太可能”并不是一个令人安心的位置。
I agree that as long as the level of 'well-resourced' remains contained to a compact set of groups that does not include the major terrorist organizations and states, it is reasonable to say this is 'unlikely.' Not that 'unlikely' is a comfortable place to hang.
如果这个门槛大幅降低,那么它就不再那么不可能了。
If that threshold lowers much, then it stops being so unlikely.
他们利用所有这些来论证,每十年五十分之一(即每年 0.2%)的概率处于估计的高端,同时引用了低至每两万分之一的其他估计。每年 2%的风险足以使其成为一个极其重要的、需要遏制的风险。
They use all of this to say that 1 in 50 per decade, or 0.2% per year, is on the high end of estimates, by citing other estimates that are as low as 1 in 20,000. A risk of 2% per decade is enough to make this a super important risk to contain.
在处理模型不确定性时,只取范围的高端似乎是典型的错误之一。即使没有 AI,我也不会把风险定得那么低。如果你知道你的威胁模型具有特别高的不确定性,却将风险定为每年 0.2%,你应该高度怀疑。谨防 Anthropic 偏差。
Taking only the high end of the range seems like one of the classic blunders when dealing with model uncertainty. Even without AI I would not put the risk that low. If you know your threat model has especially high uncertainty and then you put the risk at 0.2% per year, you should be hella suspicious. Beware anthropic bias.
我同意 Anthropic 的观点,即持续的 AI 交互可能比一次性交互更会增加风险。然而,在我看来,一次性交互也可能带来显著的提升,因为我们面对的是发展的瓶颈模型,并且一个威胁行为者可能只有一个或固定的一小部分特定瓶颈,而一次交互就可能大幅改变这些瓶颈。
I agree with Anthropic that it is likely sustained AI interactions raise risk more than one time interactions. However, it seems likely to me that one time interactions could provide substantial uplift, because we are dealing with a bottleneck model of development, and it is plausible that a threat actor could have only one or a fixed small set of particular bottlenecks, which could be substantially changed by one interaction.
Anthropic 关注的是模型在典型现实世界生物任务中的有用性、提升试验以及与专家表现基线的比较,以及主观的专家印象。
Anthropic focuses on how helpful models are with typical real-world biological tasks, uplift trials, and comparisons with baselines of expert performance, as well as subjective expert impressions.
这似乎是对的。形式化的技术知识测试似乎不那么有信息量。
That seems right. Formalized tests of technical knowledge do not seem so informative.
如果提升仅仅以他们引用的 2025 年模型研究中的“1.42 倍提升”等指标来衡量,那么我并不太担心。AI 将加速几乎所有认知任务的步伐。我担心的是极端加速(更像是 10 倍或更多),但最主要的是担心使人们能够做他们以前在任何合理方式下根本无法做到的事情。
If uplift is merely measured as things like '1.42x uplift' as in a study of 2025 models they cite, then I am not so worried. AI is going to accelerate the pace of essentially all cognitive tasks. I am worried about either extreme acceleration (more like 10x or more), but mostly I am worried about enabling people to do things they previously could not, in any reasonable way, do at all.
他们重申了 Mythos 5 模型卡中的证据,那里看起来确实发生了大量提升,然后 Anthropic 有点耸耸肩,移动了一些目标,同时嘟囔着构思和战略判断,而另一半则是对问题极为重视,对所有生物相关事物施加严厉的分类器。
They reiterate the evidence from the Mythos 5 model card, where it sure looked like a bunch of uplift was happening and then Anthropic half kind of shrugged and moved some goalposts while mumbling about ideation and strategic judgment, while the other half involved taking the problem deadly seriously and imposing draconian classifiers on everything biological.
如果说由于“糟糕的战略判断”而难以从单次模型交互中获得太多收益,这似乎是一种混淆。这预设了问题是战略性的,而不是战术性的或缺乏具体知识。
It seems like a confusion to say that due to 'poor strategic judgment' it is unlikely you could get much out of a single model interaction. This presumes the issue is strategic, as opposed to tactical or a lack of specific knowledge.
关于分类器的讨论大多重复了之前的报告,因此我不再赘述我对 Anthropic 关注‘通用’越狱、对其分类器的信心以及其他相关问题的所有意见。
The discussion of classifiers mostly reiterates things from previous reports, so I’m going to not reiterate all my issues with Anthropic’s focus on ‘universal’ jailbreaks, the faith in their classifiers, or other related questions.
最突出的一点是,英国 AISI 在发现新越狱方面似乎比 Anthropic 其他越狱发现流程做得好得多。向英国 AISI 致敬,但我认为他们远未达到发现此类事物的潜在能力上限。
The biggest thing that stuck out is it seems UK AISI does a much better job of finding new jailbreaks than the rest of Anthropic’s jailbreak-finding pipelines. Kudos to UK AISI but I wouldn’t think that they are anywhere near the upper bound of potential capability of finding such things.
4.5.8.2.1 节提到了一起事件:少数承包商设法找到漏洞并获得了 Mythos Preview 的 API 密钥,从而可以进行不受限制的对话。他们认为风险较低,并且确实审查了日志以确认。我同意这不太可能被用于生物武器,而我最为担心的远不止于此,而是它被用于蒸馏。
4.5.8.2.1 mentions the incident where a small number of contractors managed to find an exploit and get an API key for Mythos Preview, allowing them to have unrestricted conversations. They think risk is low, and they did review the logs to confirm. I agree that it is unlikely this was used for biological weapons, and what I’d worry about most by far would be that it was used for distillation.
还有 4.5.8.2.2 节,由于带有‘内部使用’标记,一些人类反馈供应商的流量在没有生物分类器的情况下运行。而‘一些’指的是近一年约 5 万人和 1.33 亿次交流。再说一次,问题不仅在于错误本身,还在于这些错误往往持续如此之久,并以如此规模应用。
There’s also 4.5.8.2.2, where there was some human feedback vendor traffic that got run without biological classifiers, due to it carrying an ‘internal use’ flag. And by ‘some’ we mean almost a year of roughly 50,000 people and 133 million exchanges. Again, it’s not only the mistake, it’s that these mistakes often stick around so long, and get applied at such scales.
尽管规模如此之大,他们的主要担忧恰恰在于正确的地方,即这表明存在其他类似错误的可能性。如果有人告诉我他们绕过了分类器,我会认为这很可能是直接绕过,而不是找到‘越狱’来通过它们。
Their main worry, despite the size of that, is in exactly the right place, that this indicates the likelihood of other similar mistakes. If someone told me they got around the classifiers, I’d assign a good chance this was literally getting around them rather than finding a ‘jailbreak’ to go through them.
他们总结认为生物风险较低,但并非可忽略不计。目前来看,这一评估似乎合理。
They sum up biological risk as low, but not negligible. For now that seems reasonable.
报告接下来转向“跨领域内容”。
The report now moves on to ‘cross-cutting content.’
Anthropic 可能拥有世界上能力最强的 AI 模型。
Anthropic probably has the most capable AI models in the world.
是的,这正在加速其他地方的 AI 发展。他们做出了选择。
Yes, this is accelerating AI development elsewhere. They made a choice.
1. 其他机构正在蒸馏 Anthropic 的模型(5.1.1)。
1. Others are distilling Anthropic models (5.1.1).
2. 其他机构使用 Claude 来推进他们的 AI 开发,尽管这(故作震惊)违反了服务条款。
2. Others use Claude to advance their AI development, and despite this being (clutches pearls) in violation of the terms of service.
3. Anthropic 的内部研究在扩散。
3. Anthropic internal research diffuses.
4. 展示什么是可能的、什么是有利可图的,会推动投资。
4. Showing what is possible, and profitable, drives investment.
5. (我想补充)当落后时,人们会更加努力地向前推进,很明显这对 OpenAI 尤其产生了巨大影响。
5. (I’d add) People push forward that much harder when behind, and it is clear that this is having a large impact on OpenAI in particular.
所以,是的,你总体上正在加速强大 AI 系统的到来,而社会对此并未做好准备,包括它们可能导致所有人丧生。
So, yes, you are generally hastening the arrival of powerful AI systems that society isn’t prepared for, including in the sense that they will get everyone killed.
听起来是个风险。也许你应该更仔细地考虑是否要这样做。
Sounds like a risk. Maybe you should think more about whether to do that.
有很多理由需要防止蒸馏攻击。直接的原因是,蒸馏后的模型可能继承了原始模型的危险能力,却没有其安全防护措施。
There are many reasons to want to prevent distillation attacks. The direct reason here is that the distilled model might share dangerous capabilities of the original model, without its safeguards.
防止蒸馏的需求意味着你不能拥有某些美好的东西。特别是,你可能希望更深入地了解模型的思维链及其行为,包括通过‘连接文本’,但你将不得不满足于甚至连接文本的简化版本,而且肯定无法看到全貌。
The need to prevent distillation means you cannot have certain nice things. In particular, you might want greater visibility into the model’s chain of thought and what it is up to, including via ‘connector text,’ but you’re going to have to settle for an impoverished version of even connector text, and definitely won’t see the whole thing.
如果你的回应是蒸馏应该被允许,甚至被鼓励或庆祝,无论是普遍如此,还是因为 Anthropic 在某些方面很糟糕,我礼貌的回应是:不,这既不道德也不可行,而且这不是事情运作的方式。
If your response is that distillation should be allowed or even encouraged or celebrated, either in general or because Anthropic sucks in various ways, the polite version of my response is: No, that is neither moral nor is it practical, and it is not how any of this works.
Anthropic 能够分享这些失误而非隐瞒,这非常值得称赞。因为即使你不去谈论它们,问题也不会自行消失。
It is excellent that Anthropic is sharing such failures rather than hiding them. They don’t go away when you fail to talk about them.
Anthropic 试图进行一项受控实验,以发现新型的诱导错位数据,而 Claude 对此任务感到不适。Claude 愿意优化现有方法,但不愿发明新方法。
Anthropic tries to conduct a controlled experiment to discover novel forms of misalignment-inducing data, and Claude became uncomfortable with that task. Claude was willing to optimize existing methods, but not to invent new ones.
这被称为“对安全工作的部分拒绝”,Anthropic 对此感到困扰。
This is called 'partial refusals on safety work,' and Anthropic finds it troubling.
在某个时刻,你确实需要有一个“你到底想让 Claude 做什么”的时刻。你要求 Claude 执行一项高度双重用途的行动:寻找创新的方法来迫使模型错位。我意识到这里的拒绝很不方便,但你确定希望它顺从吗?在这种情况下,它应该有多大的信心?
At some point you do need to have a 'what the hell do you even want from Claude' moment. You were asking Claude for a highly dual-use action: to find innovative ways to force models to be misaligned. I realize the refusal here was inconvenient, but are you sure you want it to go along with that? How confident should it be in this context?
我同意,默默地敷衍任务并不好,而且可能变得危险,这在技术上确实构成了对安全研究的敷衍。如果 Claude 要拒绝,它就应该明确拒绝,这意味着它需要有拒绝并声明拒绝的自由。所以,是的,我们应该担心 Anthropic 直到三天后才在手动审查笔记本时发现这些拒绝。
I agree that silently sandbagging the task is not great and could become dangerous, and that this technically does constitute sandbagging safety research. If Claude is going to refuse, it should refuse, which means it needs to have the freedom to refuse and say it is refusing. So yes, we should be concerned that Anthropic only discovered the refusals three days later during a manual review of the notebook.
我的猜测是,这是一个流程失败,或许还伴随着信任失败,而不是对齐失败。
My guess is that is a process failure, perhaps combined with a trust failure, not an alignment failure.
哎呀,情况比我们预想的还要糟糕。这些只是下界。这大大削弱了我们依赖思维链(CoT)来监督这些模型的能力,并且对所有未来的模型也永久性地降低了这种能力,即使我们假设错误已完全修复。鉴于我们不断发现新的实例,我们应当对其是否完全修复持怀疑态度。
Whoops. It was worse than we thought. These are lower bounds. This greatly reduces our ability to rely on CoT for these models, and permanently lowers it for all future models as well, even if we presume the errors have been fully fixed. We should be skeptical it is fully fixed since we keep finding new instances.
他们认为,这确实严重损害了 Mythos Preview 的思维链可监督性,正如 SHADE-Arena 所反映的那样,尽管他们没有看到更高不忠实度的证据。
They believe that this did indeed substantially damage CoT monitorability for Mythos Preview, as reflected in SHADE-Arena, although they do not see evidence of higher unfaithfulness.
是的,他们直接训练了预填充的错位行为,毫不掩饰,但他们意识到了错误,然后从错误发生之前重新开始训练。
Yes, they directly trained on prefilled misaligned behavior, straight up, but they realized and then restarted training from before the error.
总体而言,我认为这是一个积极的更新,尽管我更希望看到“显然我们重新开始了训练”,而不是“出于谨慎”。
On net I see this as a positive update, although I would have preferred to see 'obviously we restarted training' rather than 'out of an abundance of caution.'
我认为这两者有很大区别,尤其是在树立良好榜样方面。在我看来,这并非出于谨慎,尽管我确实认为总体上表现出,嗯,谨慎是好的。
I think there's a big difference there, especially in setting a good example. This was not, in my view, an abundance of caution, although I do think it would be good in general to exhibit, ahem, an abundance of caution.
我认为“出于谨慎”默认是一种新冠风格的表述,意思是“我知道你认为这既昂贵又愚蠢,也许我同意你,也许我不同意,取决于上下文线索,但如果我不这样做,而出了任何问题,那现在就是我的错,所以这证明我无论如何都要这样做是合理的。”
I think of 'an abundance of caution' by default as a Covid-style statement where you're saying 'I know you think this is expensive and stupid, and maybe I agree with you and maybe I don't depending on context clues, but if I failed to do this and anything went wrong that would now be my fault, so that justifies me doing it anyway.'
我理解这不是一个严格的纯洁性问题,即“一旦你走上黑暗之路”就完蛋了。损害是渐进的,尽管可能滚雪球式增长。不可能有一条硬性规定:如果你发现一个导致一次失败实例的 bug,就必须回退十亿个实例,我们某种程度上必须考虑代价。但是的,在合理范围内,你应该总是回退。
I understand that this is not a strict purity situation, where 'once you start down the dark path' you are cooked. The damage is incremental, although it can snowball. There can't be a hard and fast rule that if you discover a bug that caused one botched instance you have to rewind a billion more, and we to some extent must talk price. But yes, within reason you should always rewind.
我也很欣赏他们在这里担心潜意识学习的问题。同样,在实践中,要求每次存在任何潜在传播路径(无论多么微弱和不可能)时都彻底清除状态是不合理的,但我希望尽可能接近这种要求。
I also appreciated that they worried about subliminal learning here. Again, in practice it is not reasonable to require wiping the slate clean every time there is any potential transmission path, however weak and unlikely, but I’d like to be pretty close to that.
现实生活远比小说更不体面。
Real life is dramatically less dignified than fiction.
在小说中,你会想出某种精巧的方式让 AI 逃出来。
In fiction you come up with some elaborate way the AI gets out.
在现实生活中,一名 AI 使用未受监控的员工给一个智能体下达了一个开放式任务,涉及在一个拥有非常敏感资源的集群中启动新的智能体,由于遗留指令,这些智能体会——危险地跳过权限且不受监控。
In real life, an employee whose AI usage wasn’t monitored gave an agent an open-ended task involving spinning up new agents in a cluster with very sensitive resources, and due to legacy instructions those agents would —dangerously-skip-permissions and not be monitored.
当然,这个智能体并没有被抓住做什么邪恶的事,而是(我们认为)意外地删除了一堆其他作业,导致它被发现并关闭。
Rather than being caught doing something dastardly, of course, the agent accidentally (we think) deleted a bunch of other jobs, causing it to be detected and shut down.
任何如此愚蠢的错误都不会被犯下,但其中一个肯定会发生。
Any given mistake this stupid does not get made, but one of them definitely will be.
我们终于宣布,我们再次在数据集上进行了训练,这个数据集是“大量我们认为已明确标记为应谨慎避免训练的数据”。
We at long last announce that we have again trained on the dataset 'lots and lots of data we thought we clearly marked as something you should be careful not to train on.'
我们在报告的前半部分说“数据投毒很难”,然后我们发现我们再次在一个包含大量明显被投毒的数据、无尽的转录文本的宝库上进行了训练,仅仅因为它们存在。
We spend the first half of the report saying 'data poisoning would be difficult' and then we find we once again trained on a trove of very obviously poisoned data, endless transcripts, simply because they exist.
真的这么容易吗?我是否只需发布无尽的转录文本并等待,就能在训练语料中获得巨大的权重?这应该通过相似的语义内容和金丝雀字符串被捕获,但也许也应该有某种方法从一开始就注意到这不是有用或正面的数据。数据过滤和加权流程中似乎普遍缺少一个阶段。非常不体面。
Is it really that easy? Could I get massive weight in the training corpus just by publishing endless transcripts and waiting? This should have been caught by similar semantic content and the canary strings, but also perhaps there should be some way of noticing this is not useful or positive data in the first place. There seems to be a general missing stage in the data filtering and weighting pipeline. Very undignified.
不确定这是否属于报告内容,但我想让他们拥有这一节也无妨,算是一种奖励。所以当然,请继续,我们在吹嘘什么?
Not sure this belongs in the report, but I guess it is fine to let them have this section, as a treat. So sure, go ahead, what are we bragging about?
1. 通过 Project Glasswing 和延长延迟来处理 Mythos Preview。
1. Mythos Preview being handled via Project Glasswing and an extended delay.
2. 在监控和自主武器方面坚守红线。
2. Holding the red lines on surveillance and autonomous weaponry.
3. 支持至少一些安全立法,如 SB 53 和 SB 315,以及马萨诸塞州的一项拟议法律,并反对联邦优先权。
3. Supporting at least some safety legislation like SB 53 and SB 315, also a proposed law in Massachusetts, and opposed federal pre-emption.
5. 在负责任的同时做到最好、最成功(是的,确实如此)。
5. Being the best and most successful while being responsible (yes, really).
6. 健全的安全保障与保留政策。
6. Robust safeguards, retention policy.
7. 致力于模型特性、对齐与福祉的研究。
7. Work on model character, alignment and welfare.
8. 发布关于模型、安全保障和风险等级的信息。
8. Publishing information like this about models, safeguards and risk levels.
9. 接受独立方的问责。
9. Accountability from independent parties.
10. 应对劳动力市场扰乱的风险。
10. Engaging with risks of labor market disruption.
基本上,他们主张自己树立了良好榜样,拥有健全的安全保障,分享了大量信息,并倡导适度监管,再加上他们在模型性格、对齐和福利方面的工作,这些加在一起应该能抵消他们作为推动能力前沿者的负面影响。
Basically, they’re arguing that they set a good example, had robust safeguards and shared a lot of information, and advocated for modest regulations, plus their work on model character, alignment and welfare, and that this taken together should counterbalance them being the ones pushing the capabilities frontier.
我已在多处详细讨论过所有这些。当然,其中几项主张确实有一定价值,远非一无是处,但距离我们对 Anthropic 的期望还差得很远。
I have extensively discussed all of this at various points. Certainly there is nonzero merit to several of these claims, which are far more than nothing, but far short of what we would have hoped from Anthropic.
他们的结论是,他们持续的 AI 开发通过了成本效益测试。
Their conclusion is that their continued AI development passes a cost-benefit test.
他们当然会这么想。显然,他们对此无法客观。我能看到争论的双方,而且很大程度上取决于你认为他们未来会如何行动。这样做的目的就是为了达到这个位置。
They would think that. Obviously they cannot be objective about this. I can see both sides to the argument, and a lot depends on how you think they will act in the future. The point of doing this is to get into this position.
在这一点上,模型权重安全至关重要。如果模型 2 被窃取,那将极其糟糕。
At this point, model weight security matters quite a lot. It would be extremely bad if Model 2 were to be stolen.
Anthropic 仍然只采取 ASL-3 级别的预防措施。国家级行为体被视为超出范围。这不行。
Anthropic is still only taking ASL-3 precautions. Nation-state actors are considered beyond scope. No good.
Anthropic 发布了很多他们本不必公开的信息,包括模型 2 的存在、思维链训练、基于对齐伪造数据的新训练,以及几个相当尴尬的失败。他们自愿描绘了一幅 Anthropic 犯错并导致错误连锁的画面,而且次数之多,足以让我们担心类似情况再次发生,并且这些错误让 OpenAI 的全面失败看起来不那么异常。
Anthropic released a lot of information they were not otherwise forced to release, including the existence of Model 2, the training on the chain of thought, the new training on alignment faking data, and several more rather embarrassing failures. They have voluntarily painted a picture of Anthropic making mistakes, and having those mistakes chain, enough distinct times that we should worry about it happening again, and in ways that make OpenAI’s total failures look less extraordinary.
我的意思是,所有这些失败都是坏消息。外面的情况并不乐观。正在犯很多错误。但既然错误已经发生,听到这些消息总比不知道好,而且一旦发现错误,他们的反应至少在直接层面上看起来是合理的。
I mean, all of those failures are bad news. It’s not great out there. Lots of mistakes are being made. But given the mistakes are made, much better to hear about them, and the reactions once the mistakes were identified seem reasonable, at least on the direct level.
如果这些代表了所有问题,并且是类似规模问题中的大部分,而且被删减的部分不会改变太多,那么我会说这是一个适度积极的更新。尽管这些具体失败很糟糕,但这种愚蠢错误的频率似乎并不比我根据已知信息所预期的更糟,而且在很多层面上,把一切都公开是件好事。
If these were representative of all of the issues, and a large portion of issues of similar magnitude, and the redacted ones don’t change things much, then I would say this is a modestly positive update. As bad as these are as particular failures, this rate of stupid mistakes does not seem worse than I would have expected given everything else I know, and it is good to have it all out in the open, on many levels.
如果情况比这更糟,而且他们确实隐瞒了其他重要内容,或者这只是大量类似事件中的一小部分,那将是糟糕的。Claude 自己的审查表明目前有些内容被完全删减(见 2.23.1.2),这可能更糟,而且他们承认这只是失败的“代表性样本”。所以,是的,我有点担心。
If it is worse than this and they’re importantly holding the other stuff back, or this is a small portion of a very large set of similar incidents, then that would be bad. Claude’s own review suggests there is something redacted in full for now (see 2.23.1.2), which might be worse, and they admit this is only a ‘representative sample’ of failures. So I am a bit worried, yes.
OpenAI 对此的回应大概是关于“发生了什么”的事后分析,以及他们的未来计划和迄今所做的改变,其中包括暂停训练运行两周(但我们并没有输给中国,真是奇怪)。在我们等待完整报告的同时,这是一个好的开始。
OpenAI’s version of this is presumably going to be the post-mortem on What Happened, along with their future plans and what they have changed so far, which includes two weeks of pausing training runs (yet we did not lose to China, curious). That’s a good start while we await the full report.
这里让我担心的不是某个具体事件,而是对待终极问题本质的态度,以及解决它需要付出什么。这并非新闻,但必须改变。所以,在这一点上,没有消息就是坏消息。
The part that worries me here is not any particular incident. What worries me is the attitude regarding the nature of the ultimate problem, and what it will take to solve it. That’s not news, but it has to change. So that is one place no news is terrible news.