Claude Mythos 5.1 与 Fable 5.1:能力评估

Claude Mythos 5.1 and Fable 5.1: Capabilities

兹维·莫绍维茨 Zvi Mowshowitz · Don't Worry About the Vase · 2026-09-05 · Don't Worry About the Vase ↗

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

本文评测了 Anthropic 的 Fable 5.1 模型,重点介绍其能力、定价变化及安全措施的减少。该模型在编程、科学研究和智能体任务上较前代有适度改进,在 Terminal-Bench-Science 和 CursorBench 等基准测试中表现显著。Anthropic 通过降低缓存读取价格来降低实际成本,并为符合条件的客户提供零数据保留。文章指出,用户反应不一,有人称赞其编程能力和个性,但也有人担忧其 token 消耗和成本。总体而言,Fable 5.1 被定位为强劲的竞争者,尽管部分基准测试出现倒退或被 GPT-6 Astra 超越。

This article reviews Anthropic's Fable 5.1 model, highlighting its capabilities, pricing changes, and reduced safeguards. The model shows modest improvements over previous versions in coding, scientific research, and agentic tasks, with notable gains in benchmarks like Terminal-Bench-Science and CursorBench. Anthropic has lowered effective costs by reducing cache read prices, and offers zero data retention for eligible customers. The article notes mixed reactions, with praise for coding and personality but concerns about token usage and cost. Overall, Fable 5.1 is positioned as a strong competitor, though some benchmarks show regressions or are outperformed by GPT-6 Astra.

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 14)

全文 · Full text(逐段中英对照)

目录 Table of Contents

3. 零数据保留与减少的安全保障。

3. Zero Data Retention and Reduced Safeguards.

8. 所有评审都已提交,结果非常好。

8. The Every Review Is In and It’s Very Good.

10. 我们的价格便宜,但仅按词元计费。

10. Our Price Is Cheap, But Only Per Token.

官方宣传 The Official Pitch

和往常一样,Anthropic 基本上就是说‘这是编号更大的新模型’。

As usual, Anthropic basically said, 'Here is the new model with the higher number.'

编码部门认为它非常擅长编码,而且我们的价格也很便宜。

The coding department thinks it is very good at coding, and also our price is cheap.

Alex Albert 强调了它通过代码生成视频的能力。他的例子是,他拍了一张地块的照片,然后让 Fable 设计一栋房子、渲染它,并生成一段电影般的漫游视频,链接中有短视频。

Alex Albert highlights its ability to generate videos through code. His example is that he took a picture of a property lot and told Fable to design a house, render it, and produce a cinematic walkthrough, a short video at the link.

他们用了一节来谈科学研究。这也是 OpenAI 强调的一个重点。

They spend a section on scientific research. That's also an OpenAI point of emphasis.

我们的价格便宜 Our Price Cheap

Anthropic 在 Fable 5.1 的定价上做了一件奇怪的事。标价与 Fable 5 相比没有变化,但实际价格更低,因为他们将缓存读取的价格从每百万 token 1 美元降至 0.25 美元。

Anthropic is doing a strange thing with Fable 5.1 pricing. The headline price is unchanged from Fable 5, but the effective price is lower because they are reducing the pricing on cache reads from $1 to $0.25 per million tokens.

这将使典型的每 token 成本降低 25%,而高度智能体式的工作成本降低“高达约 45%”,这与 Boris 所说的典型 Claude Code 使用成本降低 38% 一致:

This will lower typical per-token costs by 25% and highly agentic work costs by 'up to approximately 45%,' consistent with Boris's 38% cheaper for typical Claude Code use:

我认为这通过与实际成本对齐来激励,但我倾向于将部分折扣放在标价中。人是简单的生物,有时你必须用他们的语言与他们交流。

I presume this aligns incentives, by lining up with actual costs, but I would have been inclined to put some of the discount in headline costs instead. People are simple creatures, you have to talk to them on their level sometimes.

零数据保留与减少防护措施 Zero Data Retention and Reduced Safeguards

Fable 5 在 Ramp 上的支出从未超过 Anthropic 美元支出的约 11%,尽管它显然是市面上最好的模型。

Fable 5 never got above about 11% of Anthropic dollar spend on Ramp, despite being the clearly best model out there.

两大反对意见是:防护措施的爆炸半径过大,阻碍了日常工作;以及公司(通常出于各种监管考虑)未能遵守 Anthropic 的数据保留政策,该政策要求记录保留 30 天。

Two of the big objections were the safeguards having a large blast radius that stopped ordinary work, and companies, often due to various regulatory concerns, failing to abide by Anthropic's data retention policy, where they required records be kept for 30 days.

那些能够遵守的个人开发者(Dude)拥有巨大的编码优势。

The individual Dude, who could abide, had huge coding edge.

Anthropic 已经听到了你们的声音,并且有时间变得更加从容,政府也比之前因所谓的“越狱”而紧张到迫使 Fable 下线时平静了许多。Anthropic 正在开发一个新系统,允许“符合条件的客户”使用 Fable 5.1,且外部数据保留为零,通过让客户自行存储数据来实现。在此之前,这些客户可以使用完全零数据保留的功能。

Anthropic has heard you, and had some time to get more comfortable, and for the government to calm down versus when it was so nervous it forced Fable offline over a supposed 'jailbreak.' Anthropic is working on a new system that will allow 'eligible customers' to use Fable 5.1 with zero outside data retention, via letting the customer store the data. Until then, full zero data retention is available for those customers.

他们大幅降低了分类器的误报率(至少降低了 60%)。

They've greatly reduced the false positive rate on the classifiers (at least 60%).

这同时带来了太多变化,包括模型改进和价格降低,但这将是一个引人入胜的自然实验。我们应当看到 Fable 5.1 在 Anthropic 生态系统中被大规模采用,远超 Fable 5 的 11%。如果没有,那么人们确实只是被标价吓退,而没有深思熟虑。

That is too many changes at once, including model improvements and lower prices, but will be a fascinating natural experiment. We should see massive adoption of Fable 5.1 within the Anthropic ecosystem, far more than the 11% for Fable 5. If we do not, then people really are purely balking at the headline price without thinking that through.

官方基准测试 Official Benchmarks

Anthropic 分享了大量基准测试结果。总体来看,从 Fable 5 或 Opus 5 到 Fable 5.1 的提升有限,且伴有少量性能回退。我列出这些数据,以便您快速把握整体情况。目前似乎没有明显的规律。

Anthropic shares a ton of benchmarks. Mostly we see modest improvement from Fable 5 or Opus 5 to Fable 5.1, with some small regressions. I’m listing them so you can get a quick gestalt. There does not seem to be a clear pattern.

Fable 5.1 发布于 GPT-6 之前,因此以下数据反映了当时的状况。

Fable 5.1 came out prior to GPT-6, so here is where things stood at that time.

我给出的所有“斜线”分数(例如 50%/70%)分别指不使用工具和使用工具时的得分。

All ‘slash line’ scores I give, e.g. 50%/70%, refer to scores without and then with tools.

所有提升或回退的数字默认是与 Opus 5 或 Mythos 5 的最佳得分进行比较。

All improvement or regression numbers by default are versus the best score of either Opus 5 or Mythos 5.

所有数字均根据我的判断四舍五入到给定的有效位数。

All numbers are rounded to the given significant figures, as per my judgment.

生命科学评估总体显示出适度改进,并在第一篇文章中有所涵盖。

Life sciences evals overall show modest improvement and are covered in post one.

Terminal Bench 4.0 上的括号是针对 Mythos 的,其他分数是针对 Fable 的。

Parenthesis on Terminal Bench 4.0 is for Mythos, other scores are for Fable.

我根据我们能够找到的信息,将 Astra 添加到了图表中。

I have added Astra to the chart, based on what we could find.

Anthropic ECI 得分为 162.0,而 Mythos 5 为 159.5,Opus 5 为 160.7,恰好位于 Mythos 时代的趋势线上。

The Anthropic ECI score is 162.0, versus 159.5 for Mythos 5 and 160.7 for Opus 5, exactly on the Mythos-era trend line.

DeepSWE v1.1 在 5 次试验中的得分为 67.4%,但没有参考框架。

DeepSWE v1.1 score was 67.4% over 5 trials, but with no frame of reference.

FrontierCode 1.1 Extended 在更高努力水平下得分更差。Anthropic 将此归因于 Fable 5.1 在更高努力水平下无法阻止自己进行额外的有益编辑,从而导致其被标记为错误。

FrontierCode 1.1 Extended scores are worse at higher effort levels. Anthropic attributes this to Fable 5.1 being unable to stop itself from making additional helpful edits at higher effort levels, which get it marked as incorrect.

FrontierSWE v2 得分为 0.57,而 Opus 5 为 0.52,Fable 5 为 0.48,Sol 为 0.32。

FrontierSWE v2 score was 0.57 versus 0.52 for Opus 5, 0.48 for Fable 5, and 0.32 for Sol.

Terminal-Bench-Science 0.1 从 24.7% 大幅跃升至 52.6%。

Terminal-Bench-Science 0.1 was a big jump to 52.6% from 24.7%.

CursorBench 3.2 最高达到 73.4%,而 Fable 5 最高为 70.5%,且总成本显著更低。Sol 最高为 67.2%。

CursorBench 3.2 maxed out at 73.4% versus a max out of 70.5% for Fable 5, at a substantially lower total cost. Sol maxes out at 67.2%.

CritPT-Corrected 从 85.5% 略微提升至 88.4%。

CritPT-Corrected slightly improved from 85.5% to 88.4%.

ArXivMath 从 91%/91% 略微提升至 91%/94%。

ArXivMath slightly improved from 91%/91% to 91%/94%.

ProgramBench 从 86.3% 略微提升至 87.6%。

ProgramBench improved slightly from 86.3% to 87.6%.

根据图表,Humanity’s Last Exam 从 57.8%/63.8% 提升至 60.9%/65%。

As per the chart, Humanity’s Last Exam improved from 57.8%/63.8% to 60.9%/65%.

他们运行了多种多智能体测试,但没有提供良好的比较基准,因此我无法判断结果的好坏。

They run various multi-agent tests, but don’t provide good points of comparison, so I can’t tell how good the results are.

Chartography 从 37%/84% 提升至 43%/86%。

Chartography improved from 37%/84% to 43%/86%.

BenchCAD Vision2Code 从 38%/67% 提升至 44%/84%。

BenchCAD Vision2Code improved from 38%/67% to 44%/84%.

OSWorld 2.0 根据图表所示,部分匹配率从 75% 提升至 78%,严格匹配率从 39% 提升至 42%。

OSWorld 2.0 improves as per the chart from 75% partial, 39% strict to 78% partial and 42% strict.

Anthropic 更宽松版本的 GDP.pdf 没有明显改进。无工具时从 83% 提升至 85%,但在有工具时,Fable 5 仍以 87% 的高分领先,而 Fable 5.1 为 85%。奇怪的是,Fable 5.1 没有从工具中获得帮助,但这是报告的结果。

Anthropic’s more lenient version of GDP.pdf does not clearly improve. Without tools improves from 83% to 85%, but Fable 5 retains the high score with tools at 87% versus 85% for Fable 5.1. It is strange that Fable 5.1 got no help from tools but that is what was reported.

OfficeQA 得分为 80%,OfficeQA Pro 为 69%,而之前 Claude 的最高分为 79% 和 67%。

OfficeQA score was 80%, OfficeQA Pro was 69%, versus previous Claude highs of 79% and 67%.

Harvey AI 的法律智能体基准测试在最大努力下取得了 19.1% 的全通过率和 90.8% 的平均标准通过率,或在保留集上分别为 16.7% 和 93.3%。这高于 Fable 5 的 16.9% 和 13.3% 的全通过率。但法律工作是 Vals 报告出现大幅回退的子任务之一,因此如果您为此使用 Fable 5.1,我建议您检查一下。

Legal Agent Benchmark from Harvey AI scored a 19.1% all-pass rate and 90.8% mean criterion-pass under max effort, or on the held out set 16.7% and 93.3%. This is up from Fable 5’s 16.9% and 13.3% all-pass rates. But legal work is one subtask where Vals reports large regression, so I would check if you’re using Fable 5.1 for this.

GDPval-AA v2 达到 1853 ELO,而 Opus 5 之前的最高纪录为 1824。

GDPval-AA v2 achieves 1853 ELO, compared to a previous high of 1824 for Opus 5.

AA-Briefcase 在模型卡中的测量显示,rubric 通过率从 57% 提升到 61%,分析质量从 1980 提升到 2025,但在演示方面出现回退(1495 对比 1572)。AA 随后对此进行了重新采样,因此下一节报告的数字有所不同。

AA-Briefcase, as measured in the model card, improves rubric pass rate from 57% to 61% and analytic quality from 1980 to 2025, but there is a regression on presentation (1495 vs. 1572). AA then resampled this, so the next section reports different numbers.

Toolathon Verified 显示回退:pass@1 为 77.8%,对比 80.6%;pass@3 为 81.5%,对比 87%。

Toolathon Verified shows a regression: 77.8% pass@1 vs. 80.6%, or 81.5% pass@3 versus 87%.

Zapier 的 AutomationBench 从 27% 提升到 31%,但请注意 Gemini 3.7 Flash 达到了 30%。

AutomationBench from Zapier improves from 27% to 31%, but note that Gemini 3.7 Flash gets 30%.

ARC-AGI-1 和 ARC-AGI-2 没有显示出明显改进,而且由于 API 错误分类请求的问题,他们没有报告 Fable 5.1 的 ARC-AGI-3 结果,而 GPT-6 Astra 声称在 ARC-AGI-3 上达到 99.9%。Astra 在官方测试工具上得分为 62.7%,成本为 2.6 万美元。Opus 5 之前得分为 30.2%。

ARC-AGI-1 and ARC-AGI-2 do not show obvious improvements, and they don’t report ARC-AGI-3 for Fable 5.1 due to a problem with the API misclassifying requests, whereas GPT-6 Astra is claiming 99.9% on ARC-AGI-3. Astra scored 62.7% on the official harness at a cost of $26k. Opus 5 previously got 30.2%.

HealthBench 相比 Opus 5 略有下降,尽管 HealthBench Professional 略有提升。

HealthBench shows a slight regression from Opus 5, although HealthBench Professional shows a small improvement.

BioMysteryBench 略有下降,为 90.3%,而之前为 91.4%,领先于 Sol 的 86.1%。

BioMysteryBench was a small regression, 90.3% versus 91.4%, ahead of Sol at 86.1%.

他人的基准测试 Other People’s Benchmarks

Epoch 的 ECI 将 Claude Fable 5.1 和 Fable 5 均评为 163,而 Astra 从 Sol 的 162 跃升至 169。这是 Astra 最有力的数据点。

Epoch's ECI places both Claude Fable 5.1 and Fable 5 at 163, while Astra jumps ahead to 169 from Sol's 162. This is the strongest data point for Astra.

Fable 5.1 在 Artificial Analysis Index 中明显领先,得分 66,而 Opus 5 为 63,Fable 5 为 62,所有非 Anthropic 模型最高为 61,其中包括 GPT-6 Astra 的 61 分,低得奇怪。这并未能准确评估 Astra 的能力,其能力显然远高于 Sol。该指数需要调整。

Fable 5.1 took a clear lead in the Artificial Analysis Index, scoring 66 versus 63 for Opus 5, 62 for Fable 5, and a high of 61 for all non-Anthropic models, including a strangely low 61 for GPT-6 Astra. That was not a good estimate of Astra's abilities, which are clearly substantially above Sol's. This index needed a tune-up.

AA 迅速重新进行了基准测试。新版本中 Claude Fable 5.1 以 57 分领先,Astra 以 55 分紧随其后,Opus 5 为 54 分,Fable 5 和 Muse Spark 1.3 为 53 分,Sol 和 Grok 4.6 为 51 分。

AA quickly redid the benchmark. The new version has Claude Fable 5.1 in the lead at 57, followed by Astra at 55, Opus 5 at 54, and then Fable 5 and Muse Spark 1.3 at 53, with Sol and Grok 4.6 at 51.

追溯性调整不免令人起疑,但这一版本更为合理。

Retroactive adjustments are more than a little suspicious, but this is more plausible.

他们通过新增两个基准测试实现了这一点:AA-Briefcase 中 Fable 以 58% 领先,Astra 得 53%,而 Sol 仅得 49%;以及原始 GDP.pdf 中 Astra 以 33% 领先,Sol 为 28%,Fable 仅为 26%。

They got there by adding two new benchmarks, AA-Briefcase where Fable leads with 58% and Astra scores 53% but Sol only scores 49%, and the original GDP.pdf, where Astra leads with 33%, Sol is at 28% and Fable struggles with 26%.

Fable 5.1 现已成为 WeirdML 的新高,达到 92.3%,比 Fable 5 提升了 0.4 个百分点。

Fable 5.1 is now the new high in WeirdML at 92.3%, a 0.4% improvement over Fable 5.

Vals 拥有多种基准测试,综合来看 Fable 5.1 以 68.8% 领先,而 Opus 为 67.2%,Astra 为 66.6%。他们将 Gemini 3.8 Flash 排在第四位,领先于 Muse Spark 1.3。Fable 5.1 还在其“RSI 指数”中占据首位。

Vals has a variety of benchmarks, putting Fable 5.1 ahead in composite at 68.8% versus 67.2% for Opus and 66.6% for Astra. They have Gemini 3.8 Flash in fourth ahead of Muse Spark 1.3. Fable 5.1 also has the top position in their 'RSI Index'.

在 Epoch 的 FrontierMath Erdos 中,Astra 是唯一解决过问题的模型,68 题中答对 2 题,而 Fable 5.1 得分为零。Astra 还在 FrontierMath Tier 4 中以 97.6% 的成绩占据主导地位,而 Fable 5.1 为 87.8%,与 Fable 5 的 88% 几乎持平。

Epoch's FrontierMath Erdos has Astra as the only model to ever solve a problem, getting 2 out of 68, whereas Fable 5.1 got zero. Astra also dominated FrontierMath Tier 4 at 97.6%, whereas Fable 5.1 was 87.8%, almost unchanged from Fable 5's 88%.

Code Arena 在截止日期前提交了结果,两个模型都令人印象深刻,但 Astra 目前领先。他们的聊天结果仍在等待中。

Code Arena came in right at the deadline, with both models impressive, but Astra now on top. Their chat results are still pending.

系统提示 The System Prompt

一如既往,解放者普林尼会帮你搞定。这些改动看起来都很小。

As always, Pliny the Liberator has got you. The changes all seem minor.

宣传语推介 The Blurb Pitches

我已经习惯了 OpenAI 和 Anthropic 的模型被第三方公司用高度通用的“宣传语推介”来评价,说新模型很好,但这种方式最多也就是高度模板化的。

I've gotten used to OpenAI and Anthropic models having highly generic 'blurb pitches' from third-party companies, saying that the new model is good, in ways that are at best highly templated.

我注意到这次的宣传语有所不同。它们谈论的是各家公司独有的具体细节。Jane Street 的 Craig Falls 谈到了交易直觉和保持可读性。Cognition 的 Walden Yan 谈到了迁移所有 Opus 流量。诸如此类,显然都是真正打动人的东西。

I notice the blurbs this time are different. They're talking about specifics unique to different companies. Jane Street's Craig Falls talks about trading intuition and remaining readable. Cognition's Walden Yan talks about migrating all Opus traffic. It goes on from there with what are clearly things that actually impressed people.

其中有一些是 AI 写的吗?是的,有一些是 AI 写的。其中有一些是模板化的吗?是的。但很大一部分,不同寻常地,两者都不是。

Are a few of them written by AI? Yes, a few of them are written by AI. Were some of them generic? Yes. But a large percentage of them, unusually, were neither of these.

各方评论已至,评价甚佳 The Every Review Is In and It’s Very Good

这还没有考虑 Astra,但即便如此,评价也是非常正面的。

This is without taking Astra into consideration, but even so, this is very positive.

他二十分钟的视频评测和总结在这里。

His twenty-minute video review and summary is here.

他基本上说它好得惊人:编程方面是怪兽,写作方面很好,不消耗大量 token,知识工作方面表现出色,且零数据保留。

He basically says it’s amazingly great: a monster at coding, good at writing, not token hungry, good at knowledge work, zero data retention.

这里令人好奇的是提到使用更少的 token,因此有‘Opus 级价格’的说法,因为我们有其他一些报告称它使用大量 token。他在这里是将 token 使用量与 Opus 进行比较,而不是与 Fable 5 比较,且思考设置未知。

What’s curious here is the mention of using fewer tokens, hence the ‘Opus-level price,’ since we have a bunch of other reports that it uses a lot of tokens. He’s comparing token use to Opus here rather than Fable 5, at an unknown thinking setting.

积极反响 Positive Reactions

Seth Lazar 抱怨说,Astra 在那些确实不应该审查的故事上审查了它的反应,而 Fable 则正确地接受了所有内容。

Seth Lazar complains that Astra is censoring its reactions to stories where it really shouldn't, whereas Fable is correctly fine with all of it.

当然,分类器仍然存在问题:

The classifiers still have issues of course:

正如上次讨论的,在某些智能体式上下文中,由于提示注入,回退到 4.8 略有风险,尽管按照这个标准,Astra 也是如此。

As discussed last time, punting back to 4.8 is slightly unsafe in some agentic contexts due to prompt injections, although by that standard so is Astra.

有人提到迷人的个性,这与我们对 Opus 4.7、4.8 和 5 在这些方面的抱怨相反(我理解这些抱怨的来源,但我并不同意,尤其喜欢 4.8 在这方面的表现)。

There is talk of a charming personality, a reversal of the complaints we've heard about Opus 4.7, 4.8 and 5 on such fronts (I see where those complaints came from, but I did not agree with them and especially liked 4.8 on this).

对于每一个报告的趋势,总会有例外:

For every trend reported there will always be an exception:

最常见的反应似乎是高兴,但并非极度欣喜。

The modal response seems to be happiness, but not extreme delight.

我们的价格便宜,但仅按 Token 计费 Our Price Cheap But Only Per Token

有些人觉得它相当消耗 Token。

There were some who found it rather token hungry.

我的理解是,API 成本大幅降低,但订阅限制没有改变。因此,如果 5.1 版本在 Token 消耗上略高,那么在 API 上可能更便宜,但会更快耗尽你的订阅额度。或者,尽管有这些情况,它可能更贵,这正是 Artificial Analysis 在相同努力水平下报告的结果。

My understanding is that API costs are reduced substantially, but that subscription limits have not changed. So if 5.1 is modestly token hungrier, it could be cheaper in the API but exhaust your subscription faster. Or it could be more expensive despite this, which is what Artificial Analysis reports, when using identical effort levels.

努力水平影响很大,会使 Token 使用量变化 11 倍。这也可能是一个因素。

Effort level matters a lot, changing token usage by 11x. That could also be a factor.

另一种解释是,这是一个喜欢做更多事情的模型。Fable 5.1 会做更多,如果允许的话,会使用更多 Token 来完成。你也可以调低努力水平。

The other explanation is, this is a model that loves doing more. Fable 5.1 will do more, and use more tokens to do it, if allowed to do so. You can also turn down effort levels.

我肯定会在这里寻找潜在的 bug,你永远不知道。

I would definitely look for a potential bug here, you never know.

或者简单地说,“我们不为此付费,我们要么是穷人,要么是十足的傻瓜。”

Or simply, 'we don't pay for that, we are either poor or utter fools.'

价格并不便宜,你确实要么每月支付 100 美元以上,要么支付更高的 API 费率,但这个价格非常值得你所获得的东西。

The price is not cheap; you do have to either pay $100+ per month or pay the higher API rates, but the price is so very worth what you get.

如果你想说你认为 Astra 比 Fable 5.1 更适合你,那么当然,我可以理解,尽管现在下结论还为时过早。如果你想说“支付 100 美元不在考虑范围内,所以只有 Opus 算数”,并且你正在读这篇文章,那么再说一次,要么你手头非常拮据,要么你就是个傻瓜。

If you want to say Astra is better for you than Fable 5.1, then sure, I can see that, although it's too early for me to form an opinion. If you want to say 'paying $100 was off the table so only Opus counts' and are reading this then, again, either you are highly strapped for cash or you are a fool.

负面反应 Negative Reactions

据我观察,几乎没有其他明显的负面反应。‘这是个好模型,先生。’这是我能找到的最负面的非比较性表述,且不涉及对算力或成本的抱怨:

There were almost no other actively negative reactions that I could see. It’s a good model, sir. This was the most negative non-comparative statement I could find that wasn’t about being token hungry or expensive:

这也算负面,尤其是与 Sol 而非 Astra 比较时。

This also passes for negative, especially with the comparison to Sol rather than Astra.

而且,确实,很多人分辨不出来。我一直是 Fable 5.1 的爱好者,但之前也是 Fable 5 的爱好者,而且我还没有尝试过那些能明显察觉差异的任务。

Also, yes, a lot of people can’t tell. I have been a Fable 5.1 enjoyer, but I was previously a Fable 5 enjoyer, and I haven’t tried to do anything where I would have clearly noticed.

大多数负面评价来自沉默或暗示,或者说 Astra 更好。

Most of the negative is by silence or implication, and saying Astra is even better.

早期低语 Early Whispers

我将暂时担任模型福利一职,以便收集更多信息,这可能会成为一种模式。你需要更多时间来做出评估,而且这不太强调速度。目前,我们有一些来自 Wittle 和 Antra 的出色笔记。

I will be holding the model welfare post for a bit to allow for further information to be gathered, which might become the pattern. You need more time to make that assessment, and it has less of a speed premium. For now, we have some excellent notes from Wittle and Antra.

武器选择 Weapon of Choice

现在还为时过早,因为大多数人使用 Astra 还不到一天,但我们看到一些迹象表明,我 Twitter 附近的人对 Astra 感到好奇并印象深刻,像往常一样,考虑到我所在的地区异常地偏向 Claude。这里几乎没有人使用除 Claude 或 GPT 以外的任何东西。

It is very early, as most people have had Astra for less than a day, but we are seeing signs that those living near my Twitter are intrigued and impressed by Astra, as usual adjusted for my region being unusually Claude-pilled. Almost no one here uses anything except Claude or GPT.

在 Fable 5.1 和 Astra 之前,Claude 在编码和总体上对 Sol 保持着大约 2:1 的优势。

Before both Fable 5.1 and Astra, Claude held a roughly 2:1 advantage over Sol, both for coding and in general.

我计划几天后再做一次投票,看看情况如何,以观察这个新热点是保持、获得势头还是逐渐消退。一切皆有可能。

I plan to do another poll in a few days, and see how things are going, to see if the New Hotness stays, gains momentum, or fades away. All things are possible.

现在比以往任何时候都更需要:如果你是一个严肃的 AI 用户,你应该至少尝试这两个模型,看看哪个最适合你,然后做出相应的决定。

Now more than ever, if you are a serious AI user, you should try at least both these models, see which one works best for you, and decide accordingly.

互动版:图/公式 + 针对本篇提问 →