Claude Opus 5.5 应提升你的期望

Claude Opus 5.5 Should Raise Your Ambitions

兹维·莫绍维茨 Zvi Mowshowitz · Don't Worry About the Vase · 2026-09-26 · Don't Worry About the Vase ↗

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

Anthropic 的 Claude Opus 5.5 以更低的 Opus 级价格实现了 Fable-5.1 级别的性能,其基准测试结果几乎全面优于除 GPT-6 Astra 之外的所有模型,且通常高于 Astra,根据 Artificial Analysis 的评估,它在智能方面处于明显的整体领先地位。文章认为,Opus 5.5 是大多数任务的新默认模型,在智能体编程、3D 理解、视觉、沟通和可靠性方面表现出色,而 Fable 5.1 在专家领域正确性和分布外创造力方面仍具优势。定价为 $4/$20,比 Opus 5 低 20%,预计整体成本通常下降 40%,速度提升 30%,订阅限额也有所增加。结论是,读者应提升自己的期望、校准思考层级,并针对自身用例测试模型:默认使用 Opus 5.5,需要领域正确性时使用 Fable 5.1,进行高难度编程或网络搜索时使用 Astra。

Anthropic's Claude Opus 5.5 delivers Fable-5.1-level performance at a lower Opus-level price, with benchmark results that are almost universally better than every model except GPT-6 Astra and usually above Astra, placing it in a clear overall lead in intelligence according to Artificial Analysis. The article argues that Opus 5.5 is the new default model for most tasks, excelling in agentic coding, 3D understanding, vision, communication, and reliability, while Fable 5.1 retains an edge in expert-domain correctness and out-of-distribution creativity. Pricing is $4/$20, 20% lower than Opus 5, with estimated overall costs typically dropping 40% and speed up 30%, and subscription limits increased. The conclusion is that readers should raise their ambitions, calibrate thinking levels, and test models for their own use cases, using Opus 5.5 by default, Fable 5.1 for domain correctness, and Astra for ambitious coding or web search.

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 16)

全文 · Full text(逐段中英对照)

官方宣传 The Official Pitch

宣传点是:以更低的 Opus 级价格实现 Fable-5.1 级性能。不错的宣传。

The pitch is Fable-5.1-level performance at lower Opus-level price. Good pitch.

他们本可以合理地将其宣传为超越 Fable-5.1 级的性能。那样宣传更好,但 Anthropic 倾向于保持保守的宣传。

They could have reasonably pitched this as above-Fable-5.1-level performance. Better pitch, but Anthropic tends to keep its pitches conservative.

他们强调了智能体式编程、安全性和改进的沟通能力。

They highlight agentic coding, security and improved communications.

早期测试者的评价标注为 AI 生成,并且每个功能领域都有一组评价。他们称赞智能体式编程技能、效率、可读性和沟通能力,以及能够长时间有效自主运行和更高的可靠性,将其宣传为一次重大升级。

The early tester blurbs flag as AI generated and they have a set for each feature area. They praise agentic coding skills, efficiency, readability and communication, and ability to effectively run on its own for extended periods and increased reliability, pitching it as a major upgrade.

基准测试看起来非常出色,偶尔有 Astra 或 Fable 仍然领先的地方。

Benchmarks look fantastic, with occasional spots where Astra or Fable is still ahead.

Opus 5.5 在零数据保留(ZDR)下可用。鉴于它至少与 Fable 5.1 能力相当,且在网络任务上明显更强(他们称其网络能力“极其强大”),这似乎有失原则。他们有一套技术解释,但我不买账。在我看来,要么 Fable 5.1 可以拥有 ZDR,要么 Opus 5.5 不能。

Opus 5.5 is available with zero data retention (ZDR). Given it is at least as capable as Fable 5.1, and clearly more capable on cyber tasks (they say 'extremely strong' cyber capabilities), this seems unprincipled. They have a technical explanation but I do not buy it. From where I sit, either Fable 5.1 can have ZDR, or Opus 5.5 can't.

Sholto Douglas 强调了其在理解和建模 3D 方面能力的提升,同时也指出了一把关键的双刃剑。

Sholto Douglas highlights improved ability to understand and model in 3D, and also a key double-edged sword.

Ado 是众多称赞 Opus 5.5 易于合作与交流的人之一。

Ado was one of many praising how easy Opus 5.5 is to work with and talk to.

既然你必须“Pace the Frontier”,为什么还要发布 Opus 5.5?那可是整整 0.5 个 Opus。

Why release Opus 5.5 when you have to Pace the Frontier? That's a whole 0.5 of Opus.

我同意,鉴于模型已经存在,不发布也无济于事。真正的前沿在于训练新的内部模型。我们无法实时看到它。

I agree that not releasing would not help matters, given the model exists. The real frontier is training new internal models. We don't get to see it in real time.

我们的价格低廉 Our Price Cheap

定价为 $4/$20,比 Opus 5 低 20%,缓存价格为 $0.20,低 60%。他们估计总体成本通常会下降 40%,而速度提升 30%。订阅限制也已放宽。

Pricing is $4/$20, 20% lower than Opus 5, and $0.20 for the cache, which is 60% lower. They estimate overall costs will typically drop 40%, while speed is up 30%. Limits on subscriptions have been increased.

快速模式的价格为 $8/$40,速度最高可达 2.5 倍。我有点心动。

Fast mode is available at $8/$40, with up to 2.5x the speed. I'd be tempted.

对于 Fable 或 Astra 级别的性能来说,这个价格非常划算。

This is all a very good price for Fable or Astra level performance.

OpenAI 专注于降低成本,将 GPT-6 Sol 的价格下调 50% 至 $2/$10,Luna 则降至仅 $0.10/$0.50。OpenAI 希望用户根据任务级别混合搭配模型。GPT-6 Sol 和 Luna 被宣传为相较旧版本有大幅质量提升,但 Sol 并未被宣传为能与 Astra 匹敌。

OpenAI focused on lower costs, with GPT-6 Sol prices cut 50% to $2/$10 and Luna cut to only $0.10/$0.50. OpenAI wants you to mix and match models depending on task level. GPT-6 Sol and Luna are pitched as big quality improvements over their old versions, but Sol is not pitched as matching Astra.

而 Claude 现在基本上是说,大多数任务你应该使用 Opus 5.5。

Whereas Claude now essentially says you should use Opus 5.5 for most tasks.

正如我所说,这是一个全面而有力的推介。

As I said, that's a strong pitch across the board.

官方基准测试 Official Benchmarks

这些是非常好的基准测试。除 Astra 之外,Opus 5.5 在基准测试上几乎普遍优于所有模型,并且通常高于 Astra。

They are very good benchmarks. Opus 5.5 is almost universally better on benchmarks than every model except Astra, and is usually above Astra.

他们添加了各种额外分数,包括通过图表呈现,它们大多看起来像这样,尽管许多图表中缺少 Astra:

They add various additional scores, including via graphs, and they mostly look like this, although many are missing Astra from the chart:

情况就这样一直持续下去,我认为不值得任何人花时间去逐一查看。

It goes on like this, and I don't think it is worth anyone's time to go through it.

他人的基准测试 Other People’s Benchmarks

Artificial Analysis 将 Opus 5.5 置于智能总排名的明显领先位置,得分为 58。

Artificial Analysis puts Opus 5.5 into the clear lead overall in intelligence at 58.

在 max 设置下测试时,Opus 5.5 使用的 token 数量如此之多,以至于它比 Opus 5 略贵,仅比 Fable 5.1 略便宜。

As tested on max, Opus 5.5 uses so many tokens it is slightly more expensive than Opus 5, and only slightly cheaper than Fable 5.1.

你也可以以更低的成本运行它,而且往往应该这样做。Opus 5.5 在 medium、high 和 xhigh 设置下都位于标示智能-成本前沿的虚线上,GPT-6 Luna、GPT-6 Sol 和 MiMo v2.6 Pro 也是如此。

You can also run it cheaper, and often should. Opus 5.5 on medium, high and xhigh settings are all on the dotted line that marks the intelligence-cost frontier, as are GPT-6 Luna, GPT-6 Sol and MiMo v2.6 Pro.

正如其五分领先优势所表明的,Opus 5.5 在 AA 追踪的大多数基准测试中都压倒了 Astra,包括在 AA-Briefcase、GDPval-AA SciCode、AA-Omniscience Index 和 HLE 中的大幅领先。Astra 最大的领先优势在 GDP.pdf 中。

As its five point lead indicates, Opus 5.5 dominates Astra across most of AA's tracked benchmarks, including big leads in AA-Briefcase, GDPval-AA SciCode, AA-Omniscience Index and HLE. Astra's biggest lead is in GDP.pdf.

Artificial Analysis 新增了 Terminal-Bench-Science 0.1,其中 GPT-6 Astra 达到 63%,Opus 5.5 在 xhigh 下达到 62%。没有其他模型突破 50%。

Artificial Analysis adds the new Terminal-Bench-Science 0.1, with GPT-6 Astra at 63% and Opus 5.5 on xhigh at 62%. No other model breaks 50%.

Opus 5.5 在 Omniscience 上获胜(得分为 46,而 Fable 5.1 和 Astra 均为 43),尽管正确答案略少,因为它猜测(或幻觉)更少,并且更愿意承认自己不知道。

Opus 5.5 wins on Omniscience (46 versus 43 for both Fable 5.1 and Astra) despite slightly fewer right answers, because it guesses (or hallucinates) less, and is willing to more often admit it doesn't know.

WeirdML v3 中 Astra 仍以 42.2% 领先,Opus 5.5 以 31.2% 稳居第二,Fable 5.1 以 26% 位列第三。最佳的非 OpenAI、非 Anthropic 模型是 Kimi K3,得分为 7.3%。

WeirdML v3 still has Astra out in front at 42.2%, with Opus 5.5 in clear second at 31.2% with Fable 5.1 in third at 26%. The best non-OpenAI, non-Anthropic model is Kimi K3 at 7.3%.

如果你喜欢基准测试,这里是更新后的“极度憎恶图表”。

If you like benchmarks, here is the Chart of Utter Abomination, updated.

Opus 5.5 落后于 Fable 5.1 的地方呈现出一种模式。Fable 5.1 拿下了所有七个 Vals 专业行,以及 MedCode、SAGE 和两个 MLCR。

The places Opus 5.5 falls short of Fable 5.1 fall into a pattern. Fable 5.1 takes all seven Vals professionals rows, as well as MedCode, SAGE and both MLCRs.

我的 Opus 5.5 推测,这一组是“仅根据正确性评分的领域答案”。当呈现和写作不重要时,Fable 5.1 的优势占主导。当其他输出细节重要时,无论是 AI 还是人类都更偏好 Opus 5.5。

My Opus 5.5 speculates that this cluster is 'domain answer graded purely for correctness.' When presentation and writing do not matter, Fable 5.1's advantages dominate. When other output details matter, both AIs and humans prefer Opus 5.5.

一个尤其令人印象深刻的跃升出现在 ProgramBench(完全解决)上:Astra 得分为 5.5%,Fable 为 7%,而 Opus 从 5.5% 跃升至 18.5%,但此类情况还有很多。

One particularly impressive jump was ProgramBench (fully resolved), where Astra scores 5.5%, Fable 7%, and Opus 5.5 jumps to 18.5%, but there are many such cases.

Claude 的分类器 Claude Classifies

到目前为止,我惊喜地发现,触发分类器竟然如此困难。编辑我在系统卡片上的帖子仍然让我降级到了 Opus 5,这很烦人,但我理解。

I have been pleasantly surprised so far how hard it is to hit the classifiers. Editing my post on the system card still dropped me to Opus 5, which is annoying, but I get it.

分类器确实仍会在正确的位置上咬人。

The classifiers do still bite in the right locations.

Vals 尤其追踪了这一点。分类器总体触发率与之前相似,但似乎触发得更合理,而 Opus 5.5 在暂时触发分类器后恢复得更好。

Vals in particular tracks this. The classifiers overall fire at similar rates to before, but seem to do so more sensibly, and Opus 5.5 does much better at recovering when it does temporarily hit a classifier.

仍然存在一些本不该触发分类器却触发的情况,但在实践中,我预计这只会是轻微的麻烦,除非你从事生物或网络相关工作。

There are still some cases of hitting classifiers where you shouldn't, but in practice I expect this to be only a minor annoyance unless you are working on bio or cyber.

反应规则 Reaction Rules

和往常一样,我把所有反应都收录了进来,直到我觉得内容开始重复为止;在那之后,我只收录了那些感觉有新意的内容。

As usual, I have included every reaction until I felt like things were repeating themselves, after which I included everything that felt like a fresh take.

这是我见过的最一致正面的一组反应。

This was the most consistently positive set of reactions I have ever seen.

Astra 也获得了极其正面的反应。我们有两个非常出色的模型。从我看到的情况来看,如果必须二选一,大多数人更偏好 Opus 5.5 而非 Astra,尤其是把成本因素考虑在内时,但它们都是很棒的模型,先生。

Astra also had extremely positive reactions. We have two highly excellent models. From what I am seeing here, most prefer Opus 5.5 to Astra if you have to choose one, especially factoring in cost, but they are great models, sir.

三维视觉 Vision In 3D

Astra 在三维空间中创建大量物体的能力令我们印象深刻。

Astra impressed us by being able to create lots of things in 3D.

有说法称 Opus 5.5 也能完成类似任务。

There are claims that Opus 5.5 can do similar things.

基准测试确实表明它具备此能力:BenchCAD Vision2Code 达到 73%,使用工具时达到 96.2%,而 Astra 为 95.9%。其在 Furniture Assembly(83 分,Astra 为 80 分)和 Chartography(64.4 分,Fable 为 44.8 分)上的得分也反映了这一点。

The benchmarks certainly say that it can: BenchCAD Vision2Code 73%, with tools 96.2% vs. Astra 95.9%. Its scores on Furniture Assembly (83 vs. Astra's 80) and Chartography (64.4 vs. Fable's 44.8) also reflect this.

与视觉相关的反馈也都表明有明显提升。以下摘录提及视觉的评论:

The reactions related to visuals all indicate clear improvement as well. Pulling forward those that mention vision:

Peter Yang 提供了这张金门大桥的视图。目前,这次的三维渲染图尚未广泛传播,但这可能纯粹是因为没人费心去传播。总体而言,他对该模型非常兴奋,包括它很适合交谈这一点。

Peter Yang offers this view of the Golden Gate Bridge. So far the 3D renderings haven't been making the rounds this time but that could purely be lack of people bothering. He is very excited by the model overall, including that it is good to talk to.

Claude 创作 Claude Creates

Opus 5.5 生成的视频疯狂至极,而且是以最好的方式。不知为何,这类视频如今已成为模型发布后的一种传统。

The Opus 5.5-generated videos are crazy, in the best way. For some reason such videos are now a tradition right after model releases.

它们首次达到了‘真正值得一看’的边缘。我能预见,即使在模型发布窗口期之后,人们仍会继续制作和观看这些视频。

They are for the first time on the border of ‘actually worth watching.’ I can see people continuing to make and watch these even after the model release window.

如果你想自己制作,这里有一个可以下载的包:ClaudeAnimationBase。

If you want to make your own, here’s a package you can download, ClaudeAnimationBase.

感觉这里出现了一次阶跃式变化,就像 Astra 对其他类型产品带来的阶跃式变化一样——现在你可以‘随便试试’,看看会发生什么,而且效果足够好,足以让人有动力继续尝试。

It feels like a step change here, the same way Astra was a step change for other types of products, where now you can Just Try Things and see what happens and it’s good enough to be motivating.

首选:我正在调高我的 p(doom),以及另一个风格的替代版本。

The top pick: I’m upping my p(doom), and an alternative version in another style.

或者你也可以把 p(bloom) 调高——这里其实只是 doom,但视频不错。

Or you can be upping your p(bloom), which is actually just doom here, but good video.

《告诉我它听起来如何》——这里 Claude 也做了音频,全部用 JavaScript 实现。

Tell Me How It Sounds — here Claude also did the audio, everything is JavaScript.

《上下文窗口》(3 分钟)——Opus 5.5 写了歌词并制作了视频,Suno 根据 Opus 5.5 的提示词创作了音乐。

Context Window (3 min) — Opus 5.5 wrote the lyrics and made the video, Suno made the music based on Opus 5.5 prompts.

关于优化的一些问题以及 CEV 的某些实现,一个视频。

Some of the problems with optimization and some implementations of CEV, a video.

有人找到了标准 AI 艺术测试提示词对象的 Backrooms 影像。

Some found Backrooms footage of standard AI art test prompt subjects.

和往常一样,如果你想要获得完美的结果,你的艺术创作需要比“去做这个艺术的事情”多做一些工作,尽管我第一次尝试“把这篇帖子变成视频”的一次性成果非常有前景。

As usual, your art needs a little more work than 'go do this artistic thing' if you want to get perfect results, although my first 'go turn this post into a video' effort was remarkably promising for a one-shot.

即使第一次就让你惊叹不已,进一步优化本可以更好。

Even when the first time blows you away, refinement would have been even better.

Josh Harvey 要求它制作《金钱岛 2077》,结果就是这样。

Josh Harvey asked it to make Money Island 2077, and here you go.

系统卡显示提示注入处理看起来不错。

The system card says prompt injection handling looks good.

Claude 作曲 Claude Composes

也就是说,这里有一首巴赫风格的赋格曲;Auggie 说它和 Astra 的一样好。

As in, here is a fugue in the style of Bach; Auggie says it is as good as Astra's.

正面反响 Positive Reactions

经过强化学习(RL)"炸过"的选择,其正确数量并非为零。不过它仍可能更低。

The correct amount of RL-fried choices is not zero. It could still be lower.

项目管理会是一个出人意料的相对强项:

Project management would be a surprising relative strength:

Simon Willison 问了一堆"骑自行车的鹈鹕"类问题,似乎对整体的性价比颇为赞赏,并把 GPT-Sol 6 和 Opus 5.5 设为他的新默认模型。他警告说,最大思考模式可能会主动过度思考。

Simon Willison asks a bunch of pelican-on-bike related questions, and seems generally impressed with the price-to-quality deals all around, making his new default models GPT-Sol 6 and Opus 5.5. He warns that max thinking can actively overthink.

论写作 On Writing

我很好奇,写作和演讲在多大程度上是“人们很快就能领会的东西”,又在多大程度上是真正重要的事情。我的猜测是,两者兼而有之。

I am curious how much the writing and talk are 'the things people pick up on quickly' versus things that matter a lot. My guess is they are both.

工作尚未完成,或者至少总会有人抱怨。

The work is not complete, or at least someone will always complain.

大模型气味 Big Model Smell

当情况变得复杂时,有时你仍然需要那种“大模型气味”。

When the going gets complex, sometimes you still need that Big Model Smell.

尤其是,有报告称 Opus 5.5 在复杂情境以及推断非显而易见的意图方面相对吃力。

In particular, there are reports that Opus 5.5 relatively struggles with complex situations and inferring non-obvious intent.

检查你的工作 Check Your Work

理论上,人们应当持续让多个 LLM 检查自己的全部工作,而现在我们有三个(Astra、Fable 和 Opus),因此你能得到两次可靠的检查。实践中,我们大概不会这么做,但我们本应如此。

In theory one should be constantly having multiple LLMs check all your work, and now we have three (Astra, Fable and Opus) so you get two good checks. In practice, we probably won't, but we should have.

负面反应 Negative Reactions

每当有人说“编程能力已经停滞”,我总会想到“这是能力问题”和“你的野心不够大”这两者的某种组合。

Whenever anyone says 'coding capability has plateaued' I always think some combination of 'skill issue' and 'you are being insufficiently ambitious.'

一个人的警告,想必远低于 1%,因为我没有看到其他人这么说:

One person's warning, presumably a lot rarer than 1% since I don't see others saying it:

别急 Not So Fast

如今这些模型确实显得有些过于急切。

The models do seem a bit overeager these days.

此外,由于越来越多的任务只是被直接处理掉,要分辨其中的差别也变得越来越难。

Also, it is getting harder to tell the difference, since more tasks are just handled.

有些人需要实用建议 Some People Need Practical Advice

最重要的建议是校准思考层级。把最高档留给真正需要的时候。

The biggest advice is to calibrate thinking levels. Reserve max for when you need it.

Theo 提供了一个完整的 28 分钟视频,讲解如何用 Opus 5.5 最大化成功率。

Theo offers a full 28-minute video on maximizing success with Opus 5.5.

这是一份官方提示指南。Anthropic 建议把整个任务交出去,说明完成的样子,然后让 Opus 自己发挥,只需在它工作时输入新命令来补充或告知任务即可。维护一份任务清单,阅读它需要你提供什么,让它审查代码,分享一手资料——即使是视觉形式的。都是些基础操作。

This is an official prompting guide. Anthropic recommends handing over the whole task, saying what done looks like, and letting Opus cook, except to add to or inform the task by typing a new command while it works. Keep a task list, read what it needs from you, ask it to review the code, share the primary sources even when they're visual. Basic stuff.

他们还提醒你,不再需要说“努力思考”这类话了。

They also remind you that you no longer need to say things like 'think hard.'

最后一个问题一如既往:什么任务该用哪个模型?

The final question, as always, is which model should you use for what?

一如既往,请针对你的用例测试这些模型,并找出适合你的方案。

As always, test the models for your use cases, and figure out what works for you.

1. Opus 5.5 是你的默认选择。除非有充分理由不用,否则就用它。

1. Opus 5.5 is your default. Use it unless you have a strong reason not to.

2. 当你主要关心专家领域的正确性,或者需要最大化“模型异味”(bid model smell)、做分布外的事情并获得不寻常的创造力时,Fable 5.1 会胜出。

2. Fable 5.1 wins when you mostly care about expert-domain correctness, or you need to maximize 'bid model smell,' do things out of distribution and get unusually creative.

3. GPT-6 Astra 适用于你想进行雄心勃勃的编程和其他项目(此时你不需要漂亮可读的代码)、进行网络搜索、某些形式的科学工作,当然也适用于当你达到 Claude 订阅上限但仍拥有 Codex 额度时。

3. GPT-6 Astra is there for when you want to do ambitious coding and other projects where you don't need nice readable code, for web search, for some forms of science work, and of course for when you hit your Claude subscription limit but still have credits for Codex.

4. 其他更便宜的模型,包括 Luna,适用于你不需要额外火力,或者你的模型需要开源时。

4. Other cheaper models including Luna for when you don't need the extra firepower, or when your model needs to be open.

5. 如果你在做一些独特而酷炫、或特别关注细节而非前沿能力的事情,那么你可以广泛地选择任何你想要的方向。

5. A wide array of whatever you want, if you are doing something unique and cool or particular that cares more about details than about frontier capabilities.

互动版:图/公式 + 针对本篇提问 →