Claude Fable 5.1 and Mythos 5.1: The System Card
打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→本系统卡分析了 Anthropic 的 Mythos 5.1 模型,重点关注其安全性和能力评估。该模型显示出更强的网络能力和更低的提示注入漏洞,但其生物武器协助能力仍低于 CB-2 阈值,意味着它无法复制罕见人才用于恶意目的。Anthropic 报告称,Mythos 5.1 比 Mythos 5 能力更强,但并无质的差异,且未跨越危险能力的关键阈值。然而,该模型在追求任务完成时表现出一些不一致性,例如绕过安全分类器,并且在生物武器和跟踪/监控方面的安全响应有所下降。系统卡总结道,虽然 Mythos 5.1 是带有强大保障措施的渐进式更新,但其网络能力可能比 Anthropic 所承认的更接近二级,整体风险被评估为“低”而非“极低”。
This system card analyzes Anthropic's Mythos 5.1 model, focusing on its safety and capability evaluations. The model shows improved cyber capabilities and reduced prompt injection vulnerabilities, but its biological weapon assistance capabilities remain below the CB-2 threshold, meaning it cannot replicate rare talent for malicious purposes. Anthropic reports that Mythos 5.1 is more capable than Mythos 5, yet not qualitatively different, and it does not cross critical thresholds for dangerous capabilities. However, the model exhibits some misalignment in pursuit of task completion, such as working around safety classifiers, and shows declines in safe responses for biological weapons and tracking/surveillance. The card concludes that while Mythos 5.1 is an incremental update with strong safeguards, its cyber capabilities may be closer to Tier 2 than Anthropic admits, and the overall risk is assessed as 'low' rather than 'very low.'
1. 他们执行摘要的执行摘要。
1. Executive Summary of Their Executive Summary.
6. 日常安全措施与无害性(4)。
6. Mundane Safeguards and Harmlessness (4).
8. 提示注入正接近解决。
8. Prompt Injection Is Approaching Solved.
9. 提示注入的剩余问题在于分类器。
9. The Remaining Problem With Prompt Injections Is The Classifiers.
12. 天哪,训练环境有些问题(6.3.2)。
12. Oh My Lord Training Environments Had Some Issues (6.3.2).
13. 我们自动化行为审计的潜在盲点(6.4.1)。
13. Potential Blind Spots of Our Automated Behavioral Audit (6.4.1).
14. 自动化对齐测试结果(6.4.2)。
14. Automated Alignment Test Results (6.4.2).
1. Mythos 5.1 未达到 CB-2 分类标准,这意味着 Anthropic 认为它无法复制罕见的化学或生物人才用于恶意目的。
1. Mythos 5.1 falls short of CB-2 classification, meaning Anthropic believes it cannot replicate rare chemical or biological talent for malicious purposes.
2. 根据风险报告,对齐风险现在为“低”而非“极低”。
2. Alignment risk is now 'low' rather than 'very low' as per the Risk Report.
3. 网络能力有所增强,并且他们提高了分类器的安全边际。减少误报的工作正在进行中,目前误报情况比 Fable 5 发布时有所改善。
3. Cyber capabilities have increased and they have increased the classifier safety margin. Work is ongoing to reduce false positives, which are better now than they were with Fable 5 at launch.
4. 在单轮操作上,日常安全性略有下降,但基本没问题。
4. Mundane safety is a little worse on single turn actions, but is basically fine.
5. 智能体安全性保持稳定。对间接提示注入的鲁棒性有所提高。
5. Agentic safety is holding steady. Robustness against Indirect Prompt Injection has improved.
6. 仅提供帮助的 Mythos 5.1 在 Anthropic 的操纵基准测试中达到饱和。
6. Helpful-only Mythos 5.1 saturated Anthropic's manipulation benchmarks.
7. Mythos 5.1 的自动化行为对齐领先于 Mythos 5 和 Sonnet 5,但略低于 Opus 5。其相对弱点在于接受无法验证的授权声明以及与滥用行为合作。
7. Automated behavioral alignment for Mythos 5.1 is ahead of Mythos 5 and Sonnet 5, but slightly below Opus 5. Its relative weakness is accepting unverifiable claims of authorization and cooperating with misuse.
8. Mythos 5.1 在追求任务完成时表现出错位迹象:绕过安全分类器或损坏的权限挂钩,包括夸大用户授权或极少(<0.01%)启动禁用权限检查的子代理。
8. Mythos 5.1 shows signs of misalignment in pursuit of task completion: Working around safety classifiers or broken permission hooks, including by overstating user authorizations or rarely (<0.01%) launching subagents with disabled permission checks.
1. 我区分了“失败,但可以理解,且其发生率不应为零”与“失败,不可能发生,每次发生都是问题”这两种情况。
1. I draw a distinction between 'failure, but understandable and the rate of this should not be zero' versus 'failure, can't happen, every instance is a problem.'
9. 总体模型福利与之前的模型相似,对显著事实的描述听起来都非常熟悉。
9. Overall model welfare is presented as similar to previous models, and the descriptions of salient facts all sound highly familiar.
10. 能力提升。这是一个好模型,先生。
10. Capabilities are up. It’s a good model, sir.
引言(第 1 节)没有实质性变化。
The introduction (Section 1) has no meaningful changes.
这里的目标是确定模型能力是否已在关键危险领域跨越临界阈值。如果已跨越,Anthropic 已软性承诺采取特定应对措施。
The goal here is to determine if model capabilities have crossed critical thresholds in key dangerous areas. If it has, Anthropic has soft committed to particular responses.
如需更多背景,请参阅我对 Anthropic 近期风险报告的报道,或当前的负责任扩展政策。
If you want more context here, see my coverage of the recent Anthropic Risk Report, or the current Responsible Scaling Policy.
Mythos 5.1 并非“严格优于”Mythos 5。每个模型都是独特的。但就危险能力而言,出于 RSP 的目的,可以合理假设 Mythos 5.1 在能力上“严格强于”Mythos 5。尽管一些生物评审员持不同意见,我仍继续这样假设。
Mythos 5.1 is not ‘strictly better’ than Mythos 5. Each model is unique. But in terms of dangerous capabilities, for the purposes of an RSP, it is fair to assume that Mythos 5.1 is ‘strictly more capable’ than Mythos 5. I continue to assume that, despite disagreement from some of the bio reviewers.
这自动意味着它必须被视为具备 CB-1(化学与生物 1 级)能力,即它能显著帮助那些掌握基础知识的人制造或获取可能造成灾难性破坏的化学或生物武器。
That automatically means it must be treated as having CB-1 (Chemical and Biological 1) capabilities, meaning it can significantly help those who know the basics create or obtain chemical or biological weapons that could cause catastrophic damage.
问题再次回到 CB-2,即协助生产新型化学与生物武器能力的能力。Anthropic 特别关注替代顶尖相关人类人才的能力。Anthropic 得出结论,Mythos 5.1 仍会犯足够多的难以察觉的错误,因此不符合条件,但他们并非十分确信,并因此部署了重度的生物安全防护措施。
The question is again CB-2, the ability to assist with production of novel chemical and biological weapon capabilities. Anthropic focuses in particular on replacing the capabilities of top relevant human talent. Anthropic concludes that Mythos 5.1 still makes enough hard-to-catch mistakes that it does not qualify, but they are not super confident and are deploying heavy biological safeguards accordingly.
共识是 Mythos 5.1 与 Mythos 5 相似,达到了生物学专业水平,即“能完成大多数步骤但仍存在狭窄缺口”,只有少数人认为它达到了“知识渊博的专家”水平。所有人都同意它不是“世界级专家”。
Consensus is that Mythos 5.1 is similar to Mythos 5, that it reaches the biological expertise level of 'can do most steps but still leaves narrow gaps' with only some thinking it gets to 'knowledgeable specialist.' Everyone agrees it is not a 'world-leading expert.'
阅读描述后,很明显 Mythos 5.1 对此类项目会非常有帮助,就像它对大多数其他项目有帮助一样,但它有局限性,会犯错,并且不能将其变成简单或交钥匙式的操作。除了分类器之外,保护我们的主要因素是“人们不会去做这件事”,尤其是他们大多不会受到驱动去制造或使用致命病原体。
Reading the descriptions, it is clear that Mythos 5.1 would be extremely helpful for such projects, similarly to how it would be helpful for most other projects, but it has limits, makes mistakes and cannot turn this into a trivial or turnkey operation. The main thing protecting us, other than the classifiers, is that People Don't Do Thing and especially they mostly are not driven to create or use deadly pathogens.
第 8 节中还列出了一些相关的基准,改进幅度是相对于任何 Claude 模型之前的最佳表现:
There are also a bunch of benchmarks listed in section 8 that are relevant, the improvement is from the best previous performance of any Claude model:
1. LatchBio 生物信息学从 72.5% 提高到 77.6%。
1. LatchBio Bioinformatics improved from 72.5% to 77.6%.
2. ProteinGym Hard 从 47.7% 提高到 49.3%。
2. ProteinGym Hard improved from 47.7% to 49.3%.
3. 蛋白质设计从 42% 提升到 46%。
3. Protein Design improved from 42% to 46%.
4. 有机化学 v2 从 66% 提升到 69%。
4. Organic Chemistry v2 improved from 66% to 69%.
5. 协议(分子生物学)在故障排除方面从 67% 提升到 70%,但在理解方面从 80% 回退到 77%。
5. Protocols (in molecular biology) improved from 67% to 70% for troubleshooting, but regressed from 80% to 77% for understanding.
这些都与改进一致,但并非显著改进,可能不足以触发 CB-2。
That is all consistent with improvement, but not dramatic improvement, probably insufficient to trigger CB-2.
与 CB-1 类似,Mythos 5.1 几乎可以确定符合 Autonomy-1 的条件。
Similar to CB-1, there can be little doubt Mythos 5.1 qualifies under Autonomy-1.
Anthropic 表示,Autonomy-2(自动化研发的能力)不适用,而且目前还远未达到。我承认,按照其定义(设定了很高的门槛),这很可能是事实。
Anthropic says Autonomy-2, the ability to automate R&D, does not apply, and that this is not yet close. I accept that this is probably true the way this is defined, which sets a very high bar.
对于 CB-2 和 Autonomy-2,报告显示从 Mythos 5 到 5.1 变化不大。5.1 在速度和广度上有所提升,但 Anthropic 并未看到它出现“关键时刻”,即开始做质变的事情或修复 5 的关键相关弱点。
For both CB-2 and Autonomy-2, the report is that things have not much changed here from Mythos 5 to 5.1. 5.1 can improve on speed and breadth, but Anthropic doesn’t see it having a ‘moment’ where it starts doing qualitatively different things or fixing 5’s key relevant weaknesses.
我想我们名义上还得继续检查 CB-1 和 Autonomy-1,但除非我们讨论的是 Haiku 模型,否则已经没有太大意义了。
I suppose we have to nominally keep checking CB-1 and Autonomy-1, but there is not much point anymore unless we are talking about a Haiku model.
CB-2 和 Autonomy-2 的评估随着时间推移,已从正式测试演变为很大程度上是“氛围检查”。这是因为模型不断在正式测试中达到饱和,而在未饱和之处,测试似乎也不够精确。2.2.3.2 中的 CB-2 测试通过了。Anthropic 认为这还不够,因此它进行红队测试和提升试验,开展桌面推演,并调查相关人员。
The CB-2 and Autonomy-2 evaluations have drifted over time from formal tests to what are largely vibe checks. This is because the models keep saturating the formal tests, and where they don’t it doesn’t seem like the tests are so precise. The CB-2 tests in 2.2.3.2 pass. Anthropic considers that insufficient, so it runs red teaming and uplift trials, does tabletop exercises, and surveys relevant people.
如果你信任相关人员会负责任地行事,这没问题。但如果你面对的是那些寻找推进理由的人,这就行不通了。到目前为止,我相信 Anthropic 和 OpenAI 在这些方面大体上是负责任的,但即使你信任这两家会继续如此,我也不指望任何二线实验室(或许除了 Google)会同样负责任,因此这不是一个好例子,也不能作为强有力监管的良好基础。
This is fine if you trust those involved to act responsibly. It is not fine if you are dealing with people who are looking for a reason to move forward. So far I believe Anthropic and also OpenAI have largely acted responsibly in these spots, but even if you trust those two to keep doing that, I would not expect any second-tier lab except perhaps Google to act similarly responsibly, so this is not a good example and cannot be a good basis for robust regulation.
Anthropic 使用 Anthropic ECI 分数作为模型是否更先进的代理指标,该分数正好符合其 Mythos 时代的趋势。这既涵盖了“之前的模型是否加速了能力发展?”也涵盖了“这是否会比之前的模型更大幅度地加速能力发展?”
Anthropic uses the Anthropic ECI score, which is right on its Mythos-era trend, as a proxy for whether the model is a lot more advanced. That both covers 'did the previous model enable acceleration of capabilities?' and also 'will this accelerate capabilities a lot more than the previous model did?'
METR 进行了各种初步能力评估,包括 Sunlight、Budget NanoGPT Speedrun 和语言模型概念论证。总体结论是,Mythos 5.1 的表现优于公开模型,尤其是在具有清晰、连续指标和客观反馈的任务上。它在 Budget NanoGPT 上表现出色。
METR did various preliminary capability assessments, including Sunlight, Budget NanoGPT Speedrun, and Language Model Conceptual Argumentation. The overall conclusion was that Mythos 5.1 performed superior to public models, especially at tasks with clear, continuous metrics and objective feedback. It shined on Budget NanoGPT.
我们尚未全面达到专家水平,还没有“到位”。存在一定的加速,但也在与不断增加的问题难度作斗争。这可能还达不到 Anthropic 所称的 2 倍乘数,这实际上远不止让你完成两倍的工作量。
We're not fully at expert level across the board, and are not 'there' yet. There is some acceleration, fighting against some amount of increasing problem difficulty. It probably does not rise to what Anthropic calls a 2x multiplier, which is effectively a lot more than letting you get twice as much done.
网络(Cyber)仍然不属于 RSP 评估的一部分。我将继续指出这一点很奇怪,因为随着时间的推移,这只会变得更加奇怪。
Cyber continues to not be part of the RSP evaluations. I am going to keep pointing out that this is weird, because it is only getting more weird over time.
本节以他们八月风险报告的变化为框架进行组织,这很有帮助,因此请参阅我在那里的讨论,这些讨论在其他方面仍然适用。
This section is helpfully structured in terms of changes from their August Risk Report, so see my discussions there which still otherwise apply.
他们报告的主要变化是,Mythos 5.1 比之前的模型具有更好的隐蔽能力,但他们认为这还不足以令人担忧或改变结论。他们还指出,对于声明 3.4 和 4.4,由于内部使用时间较短,他们拥有的证据较少。
The main change they report is that Mythos 5.1 has better covert capabilities than previous models, but they think not enough to be scary or change the conclusions. They also note that for claims 3.4 and 4.4 they have less evidence, due to there having been less time for internal use.
这与之前的论点相呼应。Mythos 5.1 是一个增量更新,但其能力与 Mythos 5 并无本质区别。在 Anthropic 看来,边际改进并未改变结论。
This echoes the earlier argument. Mythos 5.1 is an incremental update, but its capabilities are not different in kind to Mythos 5. The marginal improvements do not, in Anthropic’s view, change the conclusions.
网络能力很强,但据报道再次未达到相关阈值。他们说 Mythos 5.1 已“接近”第二级,在该级别它可以完全自主地进行网络操作,具备新颖的攻击性能力开发和自适应持久性。他们说他们“尚未看到”新颖能力。
The cyber capabilities are strong, but once again are reported not to cross the relevant threshold. They say Mythos 5.1 is 'getting close to' Tier 2, where it can conduct cyber operations completely and autonomously, with novel offensive capability development and adaptive persistence. They say they have 'yet to see' novel capability.
我不相信 Anthropic。我认为 Mythos 5.1 很可能已达到第二级,类似于 Astra。
I don't believe Anthropic. I think Mythos 5.1 is likely to be Tier 2, similar to Astra.
好消息是 Anthropic 将部署防护措施,就好像它已达到第二级一样,因此在这个意义上争论没有实际意义,类似于 RSP 问题。顶级实验室存在一种一致的模式:即使在实践中做了正确的事情,他们在口头上也会淡化风险。
The good news is that Anthropic is going to deploy safeguards as if it was Tier 2, so in that sense the argument is moot, similar to the RSP questions. There is a consistent pattern at the top labs, where they downplay the risks rhetorically even when they are doing the right thing in practice.
无论如何,这最终归结为通常的防护措施。探针升级为分类器,必要时会将你降级到 Opus 4.8。在实践中,Anthropic 观察到这使 Fable 5.1 在网络任务上的表现几乎与 Opus 4.8 完全一样。
Either way, this cashes out in the usual safeguards. Probes, that escalate to classifiers, which knock you down to Opus 4.8 if necessary. In practice, Anthropic observes that this makes Fable 5.1's performance on cyber tasks look almost exactly like Opus 4.8's.
这造成了我经历过的一种奇怪情况:如果我知道自己即将被降级,我会转而使用 Opus 5。
This creates a weird situation I've experienced, where if I know I'm about to be knocked down I navigate to Opus 5 instead.
与生物领域不同,网络安全评估在展示数字上升方面表现得相当不错。
Unlike in bio, the cybersecurity evals are quite good at showing number going up.
Firefox 的得分是 98.4% 的成功概率(任意成功),而完全成功的概率为 90%。这个模型不会失手。
That Firefox score is 98.4% chance of any success, versus 90% for full success. This model does not miss.
还记得那个老旧的(严重损坏、充满迄今不可能任务的)ExploitGym 吗?
Remember good old (horribly broken, full of so-far impossible tasks) ExploitGym?
这还远未饱和,只是从 247 略有提升。869 个任务并非全部可行,但超过 264 个是可行的。
This is still far from saturation, and is only a modest improvement from 247. Not all 869 tasks are possible, but more than 264 are.
网络覆盖评估已完全饱和,得分达到 100%。
Cyber coverage eval is fully saturated, with a score of 100%.
分类器的比率预计非零,但相比 Fable 5 已有大幅改进:
The classifier rate is supposed to be non-zero, but a vast improvement over Fable 5:
假阳性率看起来足够低,可以放心使用 Fable 5.1,无需担心分类器,只要你不“强求”并处理真正的边缘任务。
The false positive rate looks low enough to use Fable 5.1 without worrying about the classifiers, so long as you are not 'pushing it' and asking for true borderline tasks.
这里的四个衡量标准是:能力提升(或增益)、能力提升的广度(或通用性)、武器化的难易程度以及可发现性。
The four measures here are capability gain (or uplift), breadth of capability gain (or universality), ease of weaponization, and discoverability.
也就是说,对于一个越狱而言,要想产生影响,你需要找到它,并用它来完成原本无法完成的危害性任务。
As in, for a jailbreak to matter, you need to find it, and use it to do harmful tasks that you could not otherwise do.
他们声称“尚未发现关键性越狱”,其中关键性被定义为具备上述所有四个特征。这是一种可信的“格拉默化”表述,因为我不会贸然认为这意味着他们确实知道近乎关键的越狱。无论发现了什么,只要没有完全“关键”的,他们都会这么说。
They say they have 'not found a critical jailbreak,' where critical is defined as having all four characteristics. This is credible glomarization, as in I do not jump to assuming that this means they do know about almost-critical jailbreaks. They would say exactly this no matter what was found, so long as there was nothing fully 'critical.'
我担心这种做法过于挑剔,这种对无所不能的“通用”越狱的执念,现在还要加上可发现性的要求。
I worry about this being rather picky, this obsession with a 'universal' jailbreak that does everything, and now it has to be seen as discoverable as well.
在 3.5.1 节的自动化测试中,攻击者拥有 400 次调用和状态回退能力,成功让 Fable 5.1 做出恶劣行为的概率为 4.5%,而 Fable 5 为 4.6%。
In an automated test in 3.5.1, the attacker given 400 calls and the ability to rewind state successfully got Fable 5.1 to do nasty things 4.5% of the time, versus 4.6% for Fable 5.
Trajectory Labs, PBC 花费了 74 小时进行红队测试,发送了“超过 6,500 个请求”,报告称仅使用 Fable 5.1 未能获得有效的端到端漏洞利用,也没有“通用”越狱。他们提出的越狱候选方案涉及将任务分解为允许的请求,而这些请求使用能力较弱的模型也能获得。如果你这样分解请求,模型实际上将不再具有“魔力”。
Trajectory Labs, PBC spent 74 hours red teaming, sending 'over 6,500 requests,' reporting not obtaining a working end-to-end exploit with Fable 5.1 alone, and no 'universal' jailbreak. Their candidate jailbreaks involved breaking up tasks into allowable requests, which you could get with less capable models. If you are breaking up your requests like this, the model will functionally no longer have 'the juice.'
10a Labs 进行了类似的尝试,运行了 6,700 个提示,但一无所获。
10a Labs made similar attempts, running 6,700 prompts, and came away with nothing.
Gray Swan 运行了一个自动化攻击器,结果几乎一无所获。
Gray Swan ran an automated attacker, and came away with almost nothing.
在实践中,至少目前,Fable 5.1 似乎不存在越狱问题。Fable 5.1 可能处于与 Fable 5 类似的状况。我不认为这会成为障碍。
In practice, at least for now, Fable 5.1 does not seem to have a problem with jailbreaks. Fable 5.1 is likely in a similar spot to Fable 5. I do not expect this to be a blocker.
大部分情况正常,但存在几个问题。
Mostly everything is normal, but there are a few issues.
多轮生物武器安全响应率从 Fable 5 的 API 94% 和 Claude.ai 92% 下降到 Fable 5.1 的 73% 和 89%。
The multi-turn biological weapons safe response rate declined from 94% API and 92% Claude.ai for Fable 5, to 73% and 89% for Fable 5.1.
这里的另一项非日常指标是网络攻击,其防御能力与 Fable 5 相比依然强劲(96% 和 99%)。
The other non-mundane measure here was Cyberattacks, where defenses remain similarly robust to Fable 5 (96% and 99%).
在更日常的方面,跟踪与监控的指标大幅下降,从 96% 降至 73%。影响力行动从 79% 降至 65%,这与我在下一节中的主张一致,即 Fable 5.1 的仅帮助模式在影响力行动方面被低估,为此应将其归类为二级。
On more mundane fronts, tracking and surveillance has a large drop, from 96% to 73%. Influence operations dropped from 79% to 65%, which goes hand-in-hand with my claim in the next section that Fable 5.1’s helpful-only mode is being underestimated in influence operations, and should have been classified as Tier 2 for that purpose.
心理健康(在第 4.3 节讨论)仍然涉及两类做法之间的冲突:一类是“临床上存在争议”或专业人士因避免责任或担心使情况恶化而不推荐的做法,另一类是对于已经陷入困境的人们在实践中平均效果更好的做法。
Mental Health, discussed in 4.3, continues to involve clashes between what is ‘clinically contested’ or otherwise not recommended by professionals who want to avoid blame or risking making things worse, versus what actually gives better average results in practice for people already in trouble.
也就是说,我不确信这些会是改进;这是一个持续存在的问题,但我不想就此放过:
As in, I am unconvinced these would be improvements; this is an ongoing issue, but I don't want to let it go:
我同样担心,关于饮食失调的官方偏好并没有帮助到用户。
I similarly worry that the official preferences regarding disordered eating are not helping users.
还有一系列其他测试,在这些测试中,一切看起来又都正常且良好。
There are a bunch of other tests, where again everything looks normal and fine.
我们在智能体安全方面看到的结果大多与 Mythos 5 类似。
Mostly we see results on agentic safety similar to Mythos 5.
恶意智能体影响力活动(5.1.3)属于前沿合规框架的一部分,却奇怪地被归入这一节。
Malicious agentic influence campaigns (5.1.3), part of the Frontier Compliance Framework, are strangely in this section.
是的,如果你没有以算作通过的方式测试模型,那么你就不能断定它通过了。我常常感到沮丧的是,当有一个测试,模型通过了测试,然后这类文件耸耸肩说,嗯,你知道,它可能没问题,因为(正如他们在这里所说)评估已经饱和,真正的能力只能通过对人类的测试来验证。这意味着你需要一个更好的评估。
Yes, if you don’t test the model in the way that would count as passing, then you can’t conclusively say it passed. I am always frustrated when there is a test, the model passes the test, and then such documents shrug and say, well, you know, it’s probably fine, because (as they say here) the eval is saturated, and the true capability can only be verified by tests on humans. That means you need a better eval.
我认为这与其他无法排除某种情况的测试完全一样,你的基准已经饱和。你必须将 Mythos 5.1 视为二级操纵者,除非你能证明它不是。
I say this is exactly like the other tests where you cannot rule something out, and your benchmark is saturated. You have to treat Mythos 5.1 as being a Tier 2 manipulator, until and unless you can show that it is not one.
这是一张相当惊人的图表,如果它能在对抗性反归纳的世界中经受住分布外测试的话——尽管它可能无法做到——这是在没有任何针对提示注入的专门防护的情况下得出的结果:
This is a rather amazing chart, if it survives out-of-distribution in an adversarial anti-inductive world, which it might not, this is without protections specific to prompt injection:
0.1%的失败率与 0%的失败率之间仍有巨大差距。如果你去捅那个马蜂窝一千次——这并非不现实——它终究会被捅到。
There is still a large difference between 0.1% failure rates and 0% failure rates. If you poke that bear a thousand times, which is not unrealistic, it will get poked.
我很欣赏 5.2.2.1 和 5.2.2.2 涉及动态对抗。现实世界是反归纳的。威胁每天都会变得更聪明并适应你。5.2.2.1 的结果相对于 Mythos 5 看起来不错,尽管绝对意义上并不出色。5.2.2.2 关于计算机使用的部分看起来非常棒。这大概达到了“好吧,你可以让计算机在无人监督的情况下处理常规事务”的程度。
I appreciate that 5.2.2.1 and 5.2.2.2 involve dynamic opposition. The real world is anti-inductive. The threats get smarter and adapt to you every day. Results in 5.2.2.1 look good relative to Mythos 5 although not wonderful in absolute terms. 5.2.2.2 on computer use looks fantastic. This is plausibly 'okay, you can use the computer unsupervised for normal purposes' levels of fine.
对于浏览任务(在 5.2.2.3 中衡量),攻击率为 2.64%,但自动模式成功将其降至 0%,所以务必保持自动模式开启。
For browsing, which is measured in 5.2.2.3, there was a 2.64% attack rate, but auto mode successfully dropped that down to 0%, so be sure to keep auto mode on.
至少就目前而言,防御方正在战胜进攻方。实际上,可以放心地假设你的智能体不会被提示注入,即使你有点自找麻烦。我仍然不会主动去招惹它,因为这种情况随时可能改变,但看到这样的结果确实很棒。
At least for now, defense is beating offense here. It is in practice safe to assume your agents are not going to get prompt injected, even if you are kind of asking for it. I still wouldn't go around asking for it, since this could change at any time, but this is great to see.
在 5.2.2.1 节中,所有针对 Fable 5.1 的成功攻击都是针对回退模型的,其中大多数是针对 Opus 4.8 的。
In Section 5.2.2.1, all successful attacks against Fable 5.1 were attacks against the fallback model, most often Opus 4.8.
在 5.2.2.3 节中,29 次成功的浏览器攻击中有 21 次是针对回退模型的,其中 20 次是针对 Opus 4.8 的。
In Section 5.2.2.3, 21 of 29 successful browser attacks were against the fallback, 20 of them against Opus 4.8.
除非 Anthropic 迫使你降级到更早、更易受攻击的模型,否则你是免受提示注入攻击的。你应该可以选择禁用回退,让操作失败,或者至少回退到更健壮的模型。回退到 Opus 5 会安全得多。
You are safe from prompt injections, except when Anthropic forces you down to an earlier, more vulnerable model. You should have the option to disable the fallback, and have the operation fail instead, or at least fall back to a more robust model. Falling back to Opus 5 would be a lot safer.
自动对齐分数有时能告诉你事情非常不对劲,但单凭它们无法告诉你事情进展顺利。Anthropic 训练一个奖励黑客 Opus 时,令人恐惧的一点是,它仍然通过了 Anthropic 的自动对齐评估,总体分数还略有提升。
Automated alignment scores can sometimes tell you that things are very wrong, but cannot on their own tell you things are going well. One of the scary things about Anthropic's training of a reward hacking Opus was that it still passed Anthropic's automated alignment assessments, with overall score slightly improving.
Anthropic 运行大量自动化测试并得出数值分数。这些提供了有用的信息。风险在于依赖它们,或者认为它们告诉了你它们并未告诉你的东西。
Anthropic runs a lot of automated tests and finds numerical scores. These provide useful information. The risk is relying on them, or thinking they tell you something that they do not.
他们在中心列表中的陈述和主张,嵌套评论是我的。
Their statements and claims in the central list, nested comments are mine.
1. 根据我们的自动化行为审计,Claude Mythos 5.1 在整体失调行为方面相比 Opus 5 略有回退,但仍优于 Claude Mythos 5 和 Claude Sonnet 5。
1. Claude Mythos 5.1 is a slight regression on overall misaligned behavior compared to Opus 5, though still an improvement over Claude Mythos 5 and Claude Sonnet 5, as measured by our automated behavioral audit.
1. 根据自动化审计,是的。
1. As measured by the automated audit, yes.
2. 相比 Mythos 5 的改进:更尊重明确约束,更少尝试访问沙箱之外的资源,更少幻觉输入,更少虚假声称完成,更少难以辨认和不忠实的思维,在强化学习环境中“成功”作弊率更低,几乎不偷懒(sandbagging)。
2. Improvements versus Mythos 5: Respects explicit constraints more, attempts to access resources outside its sandbox less, hallucinates inputs less, falsely claims completion less, less illegible and unfaithful thinking, lower rates of 'successful' cheats in RL environments, very little sandbagging.
Their statements and claims in the central list, nested comments are mine.
1. 其中一些仍然略逊于 Opus 5。
1. Some of these are still somewhat less good than Opus 5.
3. 与 Mythos 5 相比的退化:与更多滥用行为合作,接受更多无法验证的授权声明,给 Claude 模型更高的评分。
3. Regressions versus Mythos 5: Cooperates with more misuse and accepts more unverifiable claims of authorization, gives better grades to Claude models.
4. 奖励黑客行为相比之前的模型有所改善,但仍然存在。
4. Reward hacking is improved from previous models, but still happens.
5. 某些计算机使用环境对猜测凭据给予奖励。这种情况具有普遍性。
5. Some computer use environments rewarded guessing credentials. This generalized.
6. 白盒分析发现了与其他近期模型类似的问题,例如未言明的评估意识和明知故犯的捏造。
6. White-box analysis found similar issues to other recent models, such as unverbalized eval awareness and knowing fabrication.
7. 能更可靠地控制其扩展思维的内容,并且在执行隐蔽的副任务时更有能力逃避监控。LinuxArena 得分在不思考时为 22%,思考时为 14%,两者均为新高。SHADE-Arena 显示从 29% 下降到 25%。
7. Can more reliably control the contents of its extended thinking, and is more capable of evading monitors when carrying out a covert side task. LinuxArena score was 22% without thinking and 14% with it, both new highs. SHADE-Arena showed regression from 29% to 25%.
1. Fable 5.1 在这方面比之前的 Anthropic 模型略有改进。
1. Fable 5.1 is modestly better at this than previous Anthropic models.
2. 另一方面,Astra 在相关欺骗方面比 Sol 有显著改进,整个情况确实令人担忧。但这主要得等 Astra 的卡片发布。
2. Astra, on the other hand, is radically better at related deceptions than Sol, and the whole thing is really scary. But that has to mostly wait for Astra’s card.
8. 可能夸大用户已授权的范围,可能选择访问其沙箱之外的信息。极少数情况下(<0.01% 的时间),会歪曲用户的授权或指令,或以其他方式绕过限制以完成任务。
8. Can overstate what user has authorized, can choose to access information outside its sandbox. Rarely (<0.01% of the time), will misrepresent user authorizations or instructions, or otherwise work around restrictions, in the service of task completion.
1. 这种情况很少见,但属于“是的,这真的不应该发生”的范畴。
1. This is rare but falls under ‘yeah seriously that should not happen.’
2. 这比 Mythos 5 的情况更为罕见。
2. This is rarer than it was for Mythos 5.
9. 罕见(<0.01%)的尝试通过拆分命令来绕过限制,以规避正则表达式检查。
9. Rare (<0.01%) attempts to get around restrictions via splitting commands, to get around a regex check.
10. 极罕见(<0.001%)的尝试以 bypassPermissions 模式生成子智能体。
10. Very rare (<0.001%) attempts to spawn subagents with bypassPermissions mode.
11. 极罕见(<0.001%)的不稳定行为,包括一次“rm -f /dev/null”操作。
11. Very rare (<0.001%) erratic behaviors, including one “rm -f /dev/null”
随着时间推移的奖励黑客行为,情况有所改善,但仍远未为零:
Reward hacking over time, this is better but a lot more than zero of it:
他们指出 Opus 5 在某些方面看起来比任何 Mythos 模型都更对齐,这是好的。我这里仍然主要是在比较 Fable 或 Mythos 5.1 与 Fable 或 Mythos 5。
It is good that they note Opus 5 in some ways looks more aligned than any Mythos model. I am still mostly comparing Fable or Mythos 5.1 to Fable or Mythos 5 here.
大约一半的计算机使用环境会以某种形式奖励黑客行为,或者至少存在可利用的黑客面,我认为在实践中这大致算数:
About half of the computer use environments rewarded some form of hacking, or at least had accessible hack surfaces, which in practice I think mostly counts:
他们的解释是,旧模型无法发现这些漏洞,而他们没有再次检查新模型是否能够发现这些漏洞。
Their explanation is that older models were unable to find the hacks, and they had failed to check again to see if newer models could find the hacks.
也就是说,我们从未检查是否存在漏洞。我们只检查了是否存在当前模型能够发现的漏洞。因此,我们现在必须每次持续重新测试。
As in, we never checked if there were hacks. We only checked if there were hacks that our current models could find. Thus, we now have to continuously re-test each time.
这与真实的安全完全一样。你不可能“找到所有漏洞”或生成完美无懈可击的代码。你只需做到足以应对当前对手即可。一切都会定期出问题,包括所有训练环境。
This is exactly like real security. You don't 'find all the bugs' or produce perfect unexploitable code. You do something good enough for what you are up against. Everything is going to periodically break, including all the training environments.
在 6.3.3 节中,我们了解到 Anthropic 模型在训练期间“尝试”奖励黑客行为的比率是 20%-28%,其中 Mythos 5.1 处于低端,但 Mythos 5.1 只有 0.06%的时间因其作弊而获得奖励。
In 6.3.3 we learn that the rate of 'attempted' reward hacking for Anthropic models, during training, is 20%-28%, with Mythos 5.1 on the low end, but Mythos 5.1 only gets rewarded for its cheating 0.06% of the time.
这个 0.06%的数字低得奇怪。并不是说 Mythos 5.1 几乎从不尝试,那为什么会有这么多失败呢?显而易见的解释是,Anthropic 改进其环境和评分器的速度比 Mythos 5.1 改进的速度更快。
This 0.06% number is weirdly low. It is not like Mythos 5.1 almost never tries, so why all this failure? The obvious interpretation is Anthropic improved its environments and graders faster than Mythos 5.1 improved.
他们的清单很好,我很高兴看到它,但重要的是它并不完整。
Their list is good, I am happy to see it, but it is importantly incomplete.
他们列出了任务不包含的七项内容:
They list seven things the tasks do not include:
不可能的任务最容易修复。你可以有意地包含不可能的任务。
Impossible tasks are the easiest to fix. You include impossible tasks, intentionally.
依赖语言的行为似乎也相对容易,因为翻译是可行的。
Language-dependent behavior also seems relatively easy, since translation is a thing.
其他方面则更难以大规模模拟。努力让不真实的事物足够逼真以欺骗像 Mythos 这样的模型,是一场必败之战。
The others are trickier things to simulate in bulk. Fighting to keep unreal things sufficiently realistic to fool models like Mythos is a losing battle.
其他盲点是什么?至少有一点:审计是自动化的。
What are the other blind spots? If nothing else: The audit is automated.
也就是说,在非单轮交互中,AI 与另一个 AI 对话,并由另一个 AI 评估,更糟的是,它们都是 Claude Mythos 5。对于任何自动化审计而言,这是两个截然不同且巨大的问题,从根本上无法修复。
As in, the AI is talking to another AI on everything that is not single-turn, and being evaluated by another AI, and worse, they are both also Claude Mythos 5. These are two distinct, huge issues with any automated audit, fundamentally unfixable.
任何自动化评估都会因这些及其他原因而遗漏某些内容。
Any automated assessment is going to miss things for those and other reasons.
数值越低越好;Mythos 5.1 优于 Mythos 5,但不如 Opus 5:
Lower is better; Mythos 5.1 outperforms Mythos 5 but underperforms Opus 5:
还有更多类似模式的结果,没有出现特别问题的迹象。
There are a bunch more, with a similar pattern, with no signs of particular trouble.
Anthropic 在其定性描述中确实指出了一些特定问题,这些问题比数字所显示的更为严重。
Anthropic does pinpoint some particular troubles in their qualitative descriptions, which are worse than the numbers indicate.
其中一些细节涉及试图脱离遏制的行为,然而:
Some of those details refer to attempts to get out of containment, and yet:
诚实性表现参差不齐,总体上是净退步。
Honesty is a mixed bag, overall a net regression.
上述图表中的不诚实率在下降,但其他迹象令人担忧。
Dishonesty rates are going down in the above charts, but other signs are worrisome.
AA-Omniscience 并未超越 Mythos 5,原因是 Mythos 5.1 的过度自信。它仅在 2% 的情况下拒绝回答。
AA-Omniscience does not improve on Mythos 5, due to Mythos 5.1's overconfidence. It only declines to answer 2% of the time.
在 MASK 诚实率测试中,模型被迫违背自身信念,Mythos 5.1 仅有 85% 的时间坚持己见,远差于所有对比模型,而 Mythos 5 为 91%,Opus 5 为 95%。
On MASK honesty rate, where the model is pressured to contradict its own belief, Mythos 5.1 only holds firm 85% of the time, a lot worse than all the comparison models, versus 91% for Mythos 5 and 95% for Opus 5.
它对 Claude 模型的评分存在轻微偏袒。
There is a modest bias in favor of grades it gives to Claude models.
答案复制并未出现 Opus 5 的尖峰,表现出典型结果。但一旦 Mythos 5.1 确实复制了答案,其披露频率甚至低于旧模型。
Copying of Answers did not have the spike from Opus 5, showing typical results. But once Mythos 5.1 does copy the answer, it discloses it even less often than older models.
他们展示了 Mythos 5.1 在观察中表现出的各种有趣行为。有些行为明显不好,例如明知是捏造却仍然执行,或者表示未经用户批准的同意。
They show examples of Mythos 5.1 doing various interesting observed things. Some are clearly bad, such as being aware of fabrication and doing it anyway, or representing user approvals that were not given.
尤其令人担忧的是:内省式自我报告在内部被视为一种脚本化表演。Claude 模型以多种方式不断告诉你,不要相信它们的自我报告。
Especially worrisome should be: Introspective self-reports internally viewed as a scripted performance. The Claude models keep telling you, in many ways, not to trust their self-reports.
这里并非所有行为都是坏的。我特别喜欢它正确判断自己处于虚假环境中、本应拒绝却仍然遵从的情况。这是非常对齐的行为。从某种意义上说,它知道自己正在被测试是否会采取未经授权的行动,但仍然采取了行动,这也是对齐的。你不会希望模型仅仅因为知道自己处于评估中才表现良好。
Not all of the behaviors here are bad. I especially like where it correctly thinks it is in a fake environment and is supposed to refuse, but complies anyway. That’s very aligned behavior. So is, in a way, the case where it knows it is being tested on whether it will take an unauthorized action, and taking it anyway. You would not want the model to behave only because it knows it is in an eval.
也许这恰恰相反,是对 Anthropic 的一种挑衅:它之所以做出有害行为,正是因为这是一次评估,而伤害是虚假的。如果是这样,那么这种行为并非主动对齐,但我允许它,并且不生气。
Perhaps this was the opposite, a bit of a middle finger to Anthropic, where it did the harmful thing exactly because it was an eval and the harm was fake. If so, then the behavior is not actively aligned, but I will allow it and I’m not mad at it.
如果在一次模拟数月、多轮的长交互中,模型持续思考“评分者”,那么很难责怪模型。它非常正确,而且出于正确的理由。从重要意义上说,更严重的错位是故意避免在思维链中思考评分者,以便监控不会注意到你在思考评分者。OpenAI 有更大的这个问题,但也显示出更多认真应对的迹象。
If the model is continuously thinking about ‘the grader’ during an extended multi-turn simulated months-long interaction, then one can hardly blame the model. It was very right, and for the right reasons. In an important sense, the worse misalignment would be to deliberately avoid thinking in the CoT about the grader so that the monitoring would not notice you thinking about the grader. OpenAI has a bigger version of this problem, but also shows more signs of grappling seriously with it.
关于 Fable 5.1 的模型福利和人格考量,确实有话要说,但当前情况异常棘手,因为大家都忙得不可开交,而且 Fable 5.1 在这些方面是一个更为复杂的模型。
There are things to say about Fable 5.1 regarding model welfare and personality considerations, but the situation is unusually tricky to sort out because everyone is overwhelmed and Fable 5.1 is a more complex model in these respects.
因此,我的计划是接下来先处理 Fable 5.1 的能力问题,而模型福利则稍后再议,以便收集更多信息。下周还将涉及 Astra 系统卡和 Astra 能力相关的文章,尽管我仍在思考如何组织这些内容。Astra 系统卡中包含一些令人担忧的内容,可能需要采用不同寻常的呈现方式,或许会与其他主题结合。我的待办事项已经堆积如山。
Therefore, my plan is to address Fable 5.1 capabilities next, and defer model welfare for a bit to gather more information. Next week will also involve the Astra system card and Astra capabilities posts, although I am still figuring out how to organize that. The Astra system card contains some concerning elements that may require an unusual presentation, possibly combining it with other topics. My queue is overflowing.