On Kimi K3: Its Capabilities And Related Discontents
打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→从《Don't Worry About the Vase》发现更多内容——一个由齿轮构成的世界。既提供快速的优质短期更新,也进行长期的世界模型构建。目前专注于每周 AI 更新。探索领域包括 AI、政策、理性、医学与生育、教育及游戏。订阅即表示你同意 Substack 的使用条款,并确认已阅读其信息收集通知和隐私政策。
Discover more from Don't Worry About the Vase A world made of gears. Doing both speed premium short term updates and long term world model building. Currently focused on weekly AI updates. Explorations include AI, policy, rationality, medicine and fertility, education and games. By subscribing, you agree to Substack's Terms of Use, and acknowledge its Information Collection Notice and Privacy Policy.
了解更多来自《别担心花瓶》的内容
Discover more from Don't Worry About the Vase
一个由齿轮构成的世界。既进行快速、优质的短期更新,也进行长期的世界模型构建。目前专注于每周 AI 更新。探索领域包括 AI、政策、理性、医学与生育、教育与游戏。
A world made of gears. Doing both speed premium short term updates and long term world model building. Currently focused on weekly AI updates. Explorations include AI, policy, rationality, medicine and fertility, education and games.
订阅即表示您同意 Substack 的使用条款,并确认其信息收集通知和隐私政策。
By subscribing, you agree to Substack's Terms of Use, and acknowledge its Information Collection Notice and Privacy Policy.
Kimi K3 是一个非常优秀的模型,基准测试表现出色。假设其权重按计划发布,纯就原始能力而言,它将成为最强的开源模型。
Kimi K3 is a very good model with excellent benchmarks. Assuming its weights are released as planned, it will become, purely in terms of raw capability, the strongest open model.
不要得意忘形。不要仅凭相对优势来评判 Kimi K3。总体而言,它落后于闭源模型前沿数月,至少四个月,我的中位数估计是六个月,其中后训练差距较小,预训练差距更大。这比之前少了几个月,但现在的每个月都更有分量。
Do not get carried away. Do not judge Kimi K3 only by its relative strengths. In aggregate, it is several months behind the closed model frontier, at least four, and my median guess is six, with the post-training closer and the pre-training farther out. This is fewer months than before, but the months are denser now.
它在一定程度上是蒸馏的产物。它在基准测试上的表现可能优于其实际性能。其所有基准测试都是在最大努力下评分的,通常使用的 token 数量远多于 Fable 或 Sol 在类似测试中使用的数量。性能看起来参差不齐。Kimi 在某些事情上会非常出色,在其他事情上则没那么出色。
It is somewhat distilled. It likely overperforms on benchmarks relative to its practical performance. All its benchmarks are scored at maximum effort, typically using far more tokens than are used in similar tests by Fable or Sol. Performance looks jagged. Kimi will be excellent at some things, less so at other things.
未来几周我们将了解更多。目前访问渠道尚不稳定,真正有机会试用 Kimi K3 的人并不多,因此我对其能力的误差线比平时更大。唉,时间不等人,我们只能继续前行。
We will know more over the coming weeks. For now, access is spotty and not many people have actually had the chance to try Kimi K3, so I have larger error bars than usual around its capabilities. Alas, time waits for no one, so we press on.
它是目前最大的开源模型,参数量为 2.8T,处于 Claude Opus 可能尺寸的上限,接近 Mythos 可能尺寸的下限,这解释了其许多优势。它速度较慢,而且似乎很消耗 token。很多正面反应源于这是一个大模型,因此它至少带有相当程度的“大模型味”,并且总体上以更慢和更昂贵的代价换取一些性能提升。这是一个很好的举措,但他们要再次这么做还需要一段时间。
It is the largest open model so far at 2.8T, on the upper end of possible sizes for Claude Opus and near the bottom of possible sizes for Mythos, which explains many of its gains. It is slow and appears hungry for tokens. A lot of the positive reactions are due to this being a big model, which thus has at least a decent amount of 'big model smell' and generally trades being slower and more expensive for some performance gains. That's a great move, but it should be a while before they can do it again.
从 Claude 蒸馏显然是其中的一部分,很可能主要来自 Fable,但显然远非全部。显然,即使只有一丁点帮助,Moonshot 也会这么做。他们发布更大模型的时机很能说明问题。
Distillation from Claude is clearly part of the story, likely largely from Fable, and is clearly nothing close to the whole story. Clearly, Moonshot would do this even if it helped only a little. The timing of their releasing a bigger model is suggestive.
它同样是一个好模型,但一旦你校正基准测试上的过度表现,再去看预期的实际性能,目前尚不清楚这与你对一个拥有 2.8T 参数的 Kimi K3 的预期有多大差异。
It is again a good model, but once you correct for the overperformance on benchmarks and look at expected practical performance, it is not clear this is so different from what you would have expected from a Kimi K3 that had 2.8T parameters.
考虑到 Kimi K3 的(初步非官方)Epoch Capabilities Index 恰好在中国趋势线上。
Consider that Kimi K3’s (preliminary unofficial) Epoch Capabilities Index is exactly on the Chinese trend line.
Kimi K3 绝对值得检查,看看它是否适合你的工作流程。在这个价格点上,无论是 API 还是订阅,它都不会取代更小、更便宜的开源模型,而且我预计它通常会输给顶尖的闭源模型,但在某些地方,Kimi K3 会是一个不错的选择。
Kimi K3 is absolutely worth checking to see if it fits into your workflows. At this price point, for both the API and the subscription, it is not going to fill the role of the smaller cheaper open models, and I expect it to usually lose out in a fight with the top closed models, but there are going to be some places where Kimi K3 is a good choice.
2. 我们曾有过那一刻(2025 年 6 月再现)。
2. We Had a Moment (Reprise from June 2025).
4. Kimi K3 公告、推介与基本事实。
4. The Kimi K3 Announcement, Pitch and Basic Facts.
8. 技术保障?那是什么?
8. Technical Safeguards? What Are Those?
11. 有些事情很难让 Kimi 去做。
11. Things It Is Not Easy to Get Kimi to Do.
12. 开放权重模型不安全,且无解。
12. Open Weight Models Are Unsafe and Nothing Can Fix This.
13. Dean Ball 试图提出建设性意见。
13. Dean Ball Attempts To Be Constructive.
14. OpenAI 员工对这一前景相对看好。
14. OpenAI Employees Are Relatively Bullish On This One.
15. Kimi K3 在典型的智能体式编码和三维生成方面相对最强。
15. Kimi K3 Is Relatively Strongest At Typical Agentic Coding and 3D.
所有关于中国模型的讨论都笼罩在 DeepSeek 时刻的阴影之下。
All discourse about Chinese models lives in the shadow of the DeepSeek moment.
有很多人非常、非常希望另一个 DeepSeek 时刻发生。
There are a lot of people who really, really want another DeepSeek moment to happen.
这些人非常、非常想讲述这样一个故事:中国的开放模型正在追赶美国的闭源模型,AI 和推理将变得商品化。
These people really, really want to tell the story that Chinese open models are catching up to American closed models, that AI and inference will become commoditized.
他们的动机各不相同。他们常常想肯定开放模型,或‘技术栈’的重要性。另一些人只是希望看到 OpenAI 和 Anthropic 倒下,或者知道这样的说法很有市场。通常最终目的是反对所有 AI 监管,或反对任何可能‘拖慢我们’或让我们‘输给中国’的事情。
Their motivations vary. They often want to affirm open models, or the importance of the ‘tech stack.’ Others simply want to see OpenAI and Anthropic go down, or know that such claims sell. Often the ultimate objective is to argue against all AI regulations, or anything that might ‘slow us down’ or cause us to ‘lose to China.’
用‘那么最好把算力卖给他们,让他们能运行这些模型,还能构建更好的模型’来回应‘现在中国有更好的模型’,简直是在主动找死。然而每次,是的,总会有人这么主张。唉。
It is actively suicidal to respond to ‘the Chinese have better models now’ with ‘then we had better sell them the compute so they can run them and also build even better ones.’ Yet every time, yes, people will argue that. Sigh.
他们往往只是想让美国 AI 停止采取预防措施,停止烦人,并移除分类器,比如‘精灵已经出了瓶子,所以释放拥有无限愿望的更大精灵吧,这是唯一的办法。’人们真的愿意承担重大灾难性风险,也不愿与分类器打交道,而且对此非常愤怒。
Often they simply want to tell American AI to stop taking precautions, to stop being annoying and take down the classifiers, as in ‘genie is out of the bottle, so release the bigger genie with unlimited wishes, it’s the only way.’ People really would take major catastrophic risks rather than deal with classifiers, and are Big Mad about this.
谷歌当日下跌 4.4%,SpaceX 下跌 3.1%,英伟达下跌超过 2%,科技股周五再次下跌,所以我们很可能会再来一次。
Google was down 4.4% on the day, SpaceX was down 3.1% and Nvidia down over 2%, and tech stocks were down again on Friday, so plausibly we’re doing this again.
是的,我们有再次这样做的风险:
And yep, we are at risk of doing this again:
2. 引用了某个令人印象深刻的基准测试。
2. There is some impressive benchmark cited.
3. 因此,美国的领先地位已经消失,证毕,就这样,不,真的,就这样。
3. Therefore, America’s lead is gone, QED, that’s it, no, really, that’s it.
在 Axios 的案例中,所涉及的基准是 Arena。仅基于这一点,他们就断言美国的领先地位已经消失。
In the Axios case, the benchmark in question is Arena. Based on that alone, they state as fact that America's lead is gone.
我本想忽略,但这种逻辑风格已多次说服华盛顿特区的许多人,并且产生了重大的政策影响。
I would ignore it, but this style of logic has convinced a lot of Washington, D.C., multiple times, and that has had substantial policy impact.
一想到这些人如今还知道了 Mythos,他们会做什么,我就不寒而栗。对 Fable 越狱的困惑很容易蔓延成更大范围的愚蠢恐慌。
I shudder to think what such folks might do now that they also know about Mythos. The confusion over Fable jailbreaks could easily extend to a broader, dumb panic.
最初的 DeepSeek 时刻之所以发生,是因为多种事件交汇在一起。
The original DeepSeek moment happened because of a confluence of events.
我们都记得 DeepSeek 时刻,它导致了 App Store 恐慌、大量股市动荡——这些动荡在基本面上几乎毫无意义,而且事后被证明相当愚蠢——以及非常紧张的一周,最终结论是终究不必恐慌。
We all remember The DeepSeek Moment, which led to Panic at the App Store, lots of stock market turmoil that made remarkably little fundamental sense and that has been borne out as rather silly, a very intense week and a conclusion to not panic after all.
几个月后,人们对(大部分)情况逐渐有了清晰的图景:多种叙事因素交汇在一起,将 DeepSeek 的 R1 从一个令人印象深刻但并不太意外、值得更新认知的模型,变成了响彻全球的一声枪响,尽管它并没有直接的“大张旗鼓”。
Over several months, a clear picture emerged of (most of) what happened: A confluence of narrative factors transformed DeepSeek’s R1 from an impressive but not terribly surprising model worth updating on into a shot heard round the world, despite the lack of direct ‘fanfare.’
尤其是以下几点共同作用,导致了这一效应:
In particular, these all worked together to cause this effect:
1. “六百万美元模型”的叙事。人们将 V3 的边际计算成本等同于 OpenAI、Anthropic 等美国实验室的总体预算。这就好比说 DeepSeek 在苹果上的花费比 OpenAI 在食物上的花费要少得多。如果做苹果对苹果的同口径比较,DeepSeek 的花费确实更少,但差距远没有那么悬殊。
1. The ‘six million dollar model’ narrative. People equated V3’s marginal compute costs with the overall budget of American labs like OpenAI and Anthropic. This is like saying DeepSeek spent a lot less on apples than OpenAI spent on food. When making an apples-to-apples comparison, DeepSeek spent less, but the difference was far less stark.
2. DeepSeek 同时发布了一款免费应用,界面设计极其简洁,且思维链(CoT)可见。DeepSeek 当时属于快速跟进者,所以没有理由隐藏 CoT。各种比较只拿 DeepSeek 的主要用例与其他地方的相同用例相比,忽视了 DeepSeek 缺乏或做得不好的功能和用例。因此,如果你想在发布首日免费查询,你得到的是一种当时独一无二且病毒式传播的体验。这迫使其他实验室也展示 CoT,并加速推出各种模型和功能。
2. DeepSeek simultaneously released an app that was free with a remarkably clean design and visible chain-of-thought (CoT). DeepSeek was fast-following, so they had no reason to hide the CoT. Comparisons only compared DeepSeek’s top use cases to the same use cases elsewhere, ignoring the features and use cases DeepSeek lacked or did poorly on. So if you wanted to do first-day free querying, you got what was at the time a unique and viral experience. This forced other labs to also show CoT and accelerate release of various models and features.
3. 要真正了解一个模型有多好需要一段时间,而不同的风格、可见的思维链(CoT)以及兴奋感让人们以为 R1 比实际更好。
3. It takes a while to know how good a model really is, and the different style, visible CoT, and excitement made people think R1 was better than it was.
4. 时机把握得无可挑剔。DeepSeek 恰在一系列其他模型发布之前入场。两周之内,局面就很清楚:美国实验室仍然领先。这是 DeepSeek 周期的顶峰,也是其他实验室周期的低谷。
4. The timing was impeccable. DeepSeek got in right before a series of other model releases. Within two weeks it was very clear that American labs remained ahead. This was the peak of a DeepSeek cycle and the low point in other labs' cycles.
5. 就技术而言,时机同样无可挑剔。当时正处于 RL Scaling(强化学习规模扩张)的极早期,训练过程仍可以低成本完成。DeepSeek 在最大化利用其芯片性能方面做得很好,但往后,它的算力劣势很可能会让它越来越吃力。
5. The timing was also impeccable in terms of the technology. This was in the very early days of RL scaling, such that the training process could still be done cheaply. DeepSeek did a great job extracting the most from its chips, but it is likely going to have increasing trouble with its compute disadvantage going forward.
6. DeepSeek 利用了‘安全测试到底算什么’这整个角度,以及快速跟进的角度,尽可能快地发布新模型——只要模型刚有一点可行性,就立刻不可撤销地放出,让它看起来比实际更领先、落后得更少。Teortaxes 指出,R1 论文列出了许多需要修复的问题,但当时 DeepSeek 没有时间修复;R1-0528 修复了这些问题,而这些修复在恐慌期间并没有被‘计入’。
6. DeepSeek leveraged the whole 'what even is safety testing' angle and the fast-following angle, shipping as quickly as possible to irrevocably release its new model the moment it was at all viable to do so, making it look relatively farther along and less behind than it actually was. Teortaxes notes that the R1 paper pointed out a bunch of things that needed fixing but that DeepSeek did not have time to fix back then, and that R1-0528 fixes them—fixes that weren't 'counted' during the panic.
7. DeepSeek 让整套‘势头’论调流行起来。此前,中国在已发布模型方面落后得多;现在 DeepSeek 的落后幅度变小了(有些人甚至说它已经领先),于是人们想:‘哦,这意味着他们很快会领先。’但并非如此,你无法做出这种假设,而且从追随者转变为领导者是一大步。
7. DeepSeek got the whole 'momentum' argument going. China had previously been much farther behind in terms of released models; DeepSeek was now less behind (and some even said it was ahead), and people thought, 'oh, that means soon they'll be ahead.' But no, you can't assume that, and moving from a follower to a leader is a big leap.
8. 这与从中国粉丝到各类对华鹰派普遍要求的“中国已赶上美国”叙事高度相关。展望未来,我们剩下的是一个“导弹差距”式的故事。
8. This was highly related to a widespread demand for a "China caught up to the USA" narrative, from China fans and also from China hawks of all sorts. Going forward, we are left with a "missile gap" style story.
9. 也有很多人一直在推动“开源模型获胜”的论点,认为非开源模型既注定失败又不算数。这些人声音很大,而且“氛围”是他们的首选武器,其中一些人与特朗普政府关系密切。
9. There are also a lot of people always pushing the "open models win" argument, and who think that non-open models are some combination of doomed and don’t count. These people are very vocal, and vibes are a weapon of choice, and some have close ties to the Trump administration.
10. 股市严重缺乏态势感知,因此他们认为这次发布比实际更重要,导致许多人“醒悟”到已知的事情,并预期其他人也会醒悟。人们对任何底层动态的运作方式都普遍存在误解,包括杰文斯悖论,以及如果你想运行 R1,就需要去买更多芯片,包括英伟达芯片。也有可能,DeepSeek 引发的股市反应在很大程度上实际上与特朗普政策公告的内幕交易有关。本质上:有效市场假说是错误的。
10. The stock market was highly lacking in situational awareness, so they considered this release much bigger news than it was, and it caused various people to "wake up" to things that were already known and anticipate others waking up, and there was widespread misunderstanding of how any of the underlying dynamics worked, including Jevons paradox and also that if you want to run R1 you go out and buy more chips, including Nvidia chips. It is also possible that a lot of the DeepSeek stock market reaction was actually about insider trading of Trump policy announcements. Essentially: The Efficient Market Hypothesis Is False.
自那时起,关于中国已经追上或正在追上的想法不断出现,在华盛顿引发了大量讨论,仿佛这一时刻的回响仍在持续。
Since then, the idea that China had caught up, or was catching up, kept coming up and drove much discussion around Washington, as echoes of this one moment.
每当中国发布一个新的最强或令人兴奋的模型,这种说法就会重新出现。中国每多一天没有发布模型,就显得落后一天;而一旦发布,就‘追上’了,看起来不再那么落后。
This is renewed every time a new strongest or exciting Chinese model comes out. Every day that China does not release a model, they look one day farther behind. When they do release, they ‘catch up’ and look less behind.
在 R1 之后、K3 之前,最有可能成为这类时刻的前 10 个候选如下:
The top 10 such potential moments since R1 and before K3 were likely these:
4. GPT-5(反向),它以一种极其愚蠢的方式吓到了很多人。
4. GPT-5 (in reverse), which spooked a lot of people in highly stupid ways.
主要是 Kimi 和 DeepSeek。Manus 掀起了一波炒作热潮,也吓到了人,而 GLM-5.2 是它们迄今为止最强的产品,让 GLM 系列崭露头角。
Mostly it’s been Kimi and DeepSeek. Manus got a hype train going and spooked people, and GLM-5.2 was by far their strongest offering, putting GLMs on the map.
其中许多都是不错的模型,但没有任何一个从根本上改变游戏格局。
Many of these were good models, but none fundamentally changed the game.
大致而言,自从 DeepSeek 时刻以来——当时 DeepSeek 大约落后八个月,但通过快速跟进在某些关键方面追赶得更快——我们一直在反复波动。有一段时间,中国似乎落后了很多。
Roughly since the DeepSeek moment, when DeepSeek was about eight months behind but had matched some key aspects faster via fast following, we have bounced around. For a while, it looked like China was quite a lot behind.
GLM-5.2 和 Kimi K3 令人印象深刻。目前对时间差距的最佳估计处于最低点。AI 领域的事件全面加速,因此尚不清楚这一差距是否比以往更少的产品周期,而且我有点预计下周还会为 Qwen 再做一次这样的评估。Kimi K3 具有 2.8T 参数,似乎仍然稳稳落后于 4 月 7 日发布的 Mythos Preview,因此这为至少三个月的差距提供了一个起点下限。
GLM-5.2 and Kimi K3 have been impressive. The current best estimate of the time gap is at its lowest point. Events in AI have accelerated all around, so it is not clear that the gap is fewer product cycles than before, and I half expect to be doing this again next week for Qwen. Kimi K3 is 2.8T and seems to still be solidly behind Mythos Preview, which was announced on April 7, so that provides a starting point lower bound of a three-month gap.
Ryan Greenblatt 尽管对 Kimi K3 感到惊喜,但根据前向传播计算估计,其预训练质量大约介于 Opus 4 和 Opus 4.5 之间,但还有其他一些优势,因此大约落后 8 个月,正如人们所预料的那样。
Ryan Greenblatt, despite being pleasantly surprised by Kimi K3, estimates that the pre-training quality is about halfway between Opus 4 and Opus 4.5 based on forward pass math, but with some other advantages, so ~8 months behind, as one might expect.
后训练的差距更小,至少部分原因是蒸馏。Moonshot 显然在创新,但显然也在蒸馏,既直接蒸馏,也通过查看输出和复制技术来快速跟进。
The post-training is closer, at least in part because of distillation. Moonshot is clearly innovating, but it is also clearly distilling, both directly and also fast following via looking at outputs and copying techniques.
英国 AISI 发布了一份涵盖 Kimi K3 之前所有情况的报告,显示狭窄网络任务的时间差距随时间推移有所缩小。完整报告见此处。
The UK AISI published a report covering everything up to Kimi K3, showing that the time gap for narrow cyber tasks has narrowed somewhat over time. Their full report is here.
[](https://substackcdn.com/image/fetch/$s_!CfQa!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6549317e-c1b1-4f7e-ab3e-21e53c183d4e_1200x913.jpeg)
[](https://substackcdn.com/image/fetch/$s_!CfQa!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6549317e-c1b1-4f7e-ab3e-21e53c183d4e_1200x913.jpeg)
我的理解是,在狭窄且相对简单的编码任务中,开放权重模型表现出相对最强的能力,而此处的基准已接近饱和。
My understanding is that narrow and relatively easy coding tasks are where open-weights models are at their relative strongest, and the benchmark here is approaching saturation.
确实,当你查看完整报告时,对于较长的任务会得到不同的答案。
Indeed, when you look at the full post, you get a different answer for longer tasks.
[](https://substackcdn.com/image/fetch/$s_!RA9C!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4f43380b-17c1-4f97-84e7-bb8c0b88dce4_3500x2160.png)
[](https://substackcdn.com/image/fetch/$s_!RA9C!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4f43380b-17c1-4f97-84e7-bb8c0b88dce4_3500x2160.png)
更长的任务在两大最值得担忧的问题上更具相关性:AI 研发自动化与网络攻击。
Longer tasks are more relevant in terms of both of the most important things to worry about: Automation of AI R&D and cyber attacks.
人们可能还会担心生物风险,尽管这种担忧已不那么流行;而且令人有些不安的是,开放模型在这方面的标准测试完全缺失。这个问题需要解决。我们确实拿到了 Andrew Ho 在 OpenAI 的 GeneBench-Pro 上的得分。该基准衡量的是计算生物学中长周期模糊性下的判断能力。Kimi K3 超出预期。Mythos 尚未接受测试。Fable 拒绝了该基准中的大多数请求。
One might also worry about bio risks, even if that is not as in fashion, and it is a little concerning we don’t see standard testing on that at all for the open models. That needs to be addressed. We do have the score on OpenAI’s GeneBench-Pro via Andrew Ho. This measures judgment under long-horizon ambiguity in computational biology. Kimi K3 exceeded expectations. Mythos has not been tested. Fable refused most requests in the benchmark.




击败 Opus 和 GPT-5.5 令人印象深刻。此前没有任何开放模型接近过这一水平。我们正处于“乱搞一通然后面对现实”的边缘。在这个领域,很可能长时间没什么事发生,直到突然间发生很多。
Beating Opus and GPT-5.5 is impressive stuff. No previous open model came close. We are on the verge of doing some f*ing around and thus finding out. This is a place where plausibly not much happens until suddenly quite a lot happens.
开放模型将在通用能力上赶上 Mythos——包括目前使其独一无二的那个特性——这一结论无可争议。这一时刻终将到来,问题只是何时到来,而这将决定实际差距的大小。考虑到 Kimi K3,我们应当预期这会在几个月后发生。
The conclusion that open models will catch up to Mythos in general capability—including the thing that currently makes it unique—is indisputable. That is coming; the question is when, and that will establish the effective gap. Given Kimi K3, we should expect this to happen a few months from now.
如果我们取 UK AISI 估算的上下限,在相对优势领域,Kimi 出现之前的差距为 4 到 7 个月,而去年为 6 到 10 个月。从绝对意义上讲,进展差距大致相当,而且一切都在加速。
If we take both ends of UK AISI's estimates, the pre-Kimi gap was 4–7 months, down from 6–10 months last year, in an area of relative strength. That is roughly a similar amount of progress gap in absolute terms, and everything is accelerating.
1. 2.8 万亿参数,896 个区域中同时激活 16 个,这意味着约 500 亿活跃参数。这不容易在本地运行,也不会那么便宜。
1. 2.8 trillion parameters, 16 of 896 areas active at once, which implies ~50B active. This is not easy to run locally, and won't be that cheap.
2. 3.00/15.00 美元,比 Opus 和 Sol 略便宜。
2. $3.00/$15.00, modestly cheaper than Opus and Sol.
3. 订阅计划为每月 19/39/99/199 美元。较大的购买有一些适度的优势,配额随价格线性增长。
3. Subscription plans are $19/$39/$99/$199 per month. The larger buys have some modest advantages and quota scales linearly with price.
6. 据报道,训练截止时间为 2026 年初。
6. Training cutoff is reportedly early 2026.
8. 所有基准测试均在最大努力设置下运行。
8. All benchmarks run under maximum effort settings.
如上所述,他们声称在官方基准测试中表现强劲,尽管不及 Sol 或 Fable。
As stated above, they claim strong official benchmarks, although short of Sol or Fable.






这则广告被泰勒·考恩称为非常正面且非常好,但在我看来却平淡无奇,并且不包含任何有用信息。
The ad, which Tyler Cowen called very positive and very good, falls flat to me, and contains zero useful information.
我们过去常看到相当露骨的刷榜行为。实验室会在测试集上训练,或者针对测试所覆盖的极窄范围进行训练,因为当时已知目标很有限。在查看数据时,你必须知道哪些实验室这么做了,以及做到什么程度。
We used to see rather explicit benchmaxxing. Labs would train on the test, or on the very narrow thing the test would cover, because we had a limited set of known targets. You had to know which labs did this, to what extent, when looking at numbers.
我们的基准测试技术已经改进,如今它们能够共同衡量真实的事物,并且有多种备用手段来防止目标设定得过于狭窄。
Our benchmarking technology has improved, and now they collectively measure real things, and there are a variety of backups in case you aim too narrowly.
整体地看待不同的基准测试也很有价值。一切都应成为一幅地图的一部分,这幅地图嵌合在一个反映底层真实状况的共性模式之中。
Looking at the gestalt of different benchmarks is also valuable. Everything should be part of a map that fits into a common pattern that reflects the underlying territory.
你完全可以继续刷榜,而不必像过去那样露骨。
You can still absolutely benchmaxx without being as explicit as you used to be.
基准测试衡量的是某些类型的能力,而非另一些;衡量的是浅层任务,而非深层任务;并且排除了许多有价值的属性或潜在的风险。此外,有些实验室比其它实验室更侧重这些方面,或者在这些方面更为成功。
Benchmarks measure some types of abilities rather than others, and measure shallow rather than deep tasks, and exclude many valuable properties or potential liabilities. And some labs focus more on those aspects, or have more success on them, than others.
你也可以将所有测试的努力程度设置为最大,Moonshot 就是这么做的。
You can also set effort to maximum for all the tests, which Moonshot did.
把基准测试视为下限。Kimi 的基准测试证明它是真实可靠的,而且它的表现最多也只能比这些基准差(或好)那么多。我过去一直预期,并且现在依然相信,这些基准略微高估了 Kimi K3 的相对能力。
Think of the benchmarks as a lower bound. Kimi’s benchmarks prove it is for real, and it could only underperform (or outperform) them by so much. I still expected, and continue to believe, that they modestly overstate Kimi K3’s relative capabilities.
在我们最接近“唯一基准”(One True Benchmark)的基准上,Kimi K3 表现优异,印证了关于该模型整体拥有第三高基准成绩的说法:
On the closest thing we have to the One True Benchmark, Kimi K3 does well, confirming claims that overall this model has the third highest benchmarks:
[](https://substackcdn.com/image/fetch/$s_!XsAE!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F19f49a9c-0198-4499-9a0e-328e51819609_761x694.png)
[](https://substackcdn.com/image/fetch/$s_!XsAE!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F19f49a9c-0198-4499-9a0e-328e51819609_761x694.png)
Kimi K3(在初步的非官方结果中)恰好位于 Epoch 能力指数(ECI)的中国趋势线上,介于 Opus 4.6 与 Opus 4.7 之间,这意味着它落后 OpenAI 和 Anthropic 六个月,但领先 Google、Meta 和 SpaceX。
Kimi K3 is (in a preliminary unofficial result) exactly on the Chinese trend line for the Epoch Capabilities Index (ECI), between Opus 4.6 and Opus 4.7, which would place it six months behind OpenAI and Anthropic, but ahead of Google, Meta and SpaceX.
[](https://substackcdn.com/image/fetch/$s_!5oKu!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe2ff303f-30c4-4cf7-98cf-630647fe3a9d_564x764.png)
[](https://substackcdn.com/image/fetch/$s_!5oKu!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe2ff303f-30c4-4cf7-98cf-630647fe3a9d_564x764.png)
Kimi K3 在 Harvey LAB-AA 全通过率上表现极为亮眼,以大幅优势位居第一。我想对此进行合理性检查:
Kimi K3 highly impresses on Harvey LAB-AA all-pass rate, in first by a wide margin, I'd like to see a sanity check on this:


就 Criterion 的通过率而言,结果为 94.6% 对 93.6%,差距虽小,但胜利仍是胜利。
For criterion pass rate this is 94.6% vs. 93.6%, less of a gap but a win is a win.
在 Arena Frontend Code 上,Kimi K3 排名第一,领先于 Fable 5 和 Sol。
Arena Frontend Code has Kimi K3 at #1 ahead of Fable 5 and Sol.
在视觉推理基准 VoxelBench 中,Kimi K3 排在 Sol 和 Fable 之后,位列第三。
Kimi K3 comes in third on VoxelBench for visual reasoning behind Sol and Fable.
明显缺失的是像 CyberGym 这样的网络安全基准测试。网络能力与编码能力并非完全脱节,因此我们可以推测。
Conspicuously missing are Cybersecurity benchmarks like CyberGym. Cyber capabilities are not so divorced from coding capabilities, so we can guess.
我找到的最接近的是 Malte Ubl 通过 DeepSec(一个私有网络基准测试)对它进行测试。结果比 Sol 低一个档次,与 GPT-5.5 类似。如果这个结果是准确的,那么网络攻击能力将有所提升,互联网在某些方面将变得更加敌对,但在其他方面会有所改善,尾部风险有限。
The closest I’ve found is Malte Ubl running it through DeepSec, a private cyber benchmark. It was a tier below Sol and similar to GPT-5.5. If that is accurate, then there will be some uplift to cyber attacks and the internet will in some ways be a more hostile place, but in other ways it will improve, and the tail risks are limited.
Parv Mahajan 报告了 CyBench 的初步结果:该基准测试在 Opus 4.7 时已经饱和,Kimi K3 也使其饱和,这只能限定性能的下界,但并没有说明太多其他信息。
Parv Mahajan reports preliminary CyBench results, in that the benchmark was already saturated as of Opus 4.7 and Kimi K3 also saturates it, which lower bounds performance but doesn’t say much else.
官方技术博客大量讨论了基准测试,也谈了一些特性,但基本上没有讨论风险或缓解措施。
The official tech blog talks a lot about benchmarks, and some about features, and talks basically not at all about risks or mitigations.
不,他们没有将 Kimi K3 提交给白宫进行 30 天审查。
No, they did not submit Kimi K3 for the 30-day review with the White House.
Moonshot 必须面对 CCP(中国共产党),这有其自身的问题。他们大概实际上不会遵守加州或欧盟的规定,而我认为加州和欧盟目前也不会对此采取任何行动。但确实,这是一个潜在的痛点,包括对任何商业使用 Kimi K3 的人所构成的风险。
Moonshot has to deal with the CCP, which comes with its own issues. They presumably will not in practice comply with California or the EU, and I presume both California and the EU will not do anything about this for now. But yes, this is one point of potential pain, including risk for anyone using Kimi K3 commercially.
Sam 就英国或其他旁观者应如何明智地看待和应对此事提出了一些好建议。随着时间的推移,我们会了解更多,包括当我们获得权重,以及像英国 AISI 这样的团队能够运行测试时。
Sam has some good advice for how the UK, or others watching, would be wise to view and react to this. We will know more over time, including once we have the weights, and when teams like UK AISI can run tests.
我认为 Lisan 对其 SVG 印象颇深,称它们优于 Fable 的 SVG,这本身就可算作一个基准。
I think it counts as a benchmark that Lisan is impressed by its SVGs, saying they are better than Fable's.
Debate Benchmark 是 Kimi K3 超出我预期的地方,而 Sol 在其中表现相对较弱。
Debate Benchmark is a benchmark where Kimi K3 outperformed my expectations, and where Sol is relatively weak.
Kimi 在 Mazur 的 Extended NYT Connections 上取得了令人印象深刻的 95.8 分,位列第三,但运行成本高于 Fable。截至上次检查,它的其他分数尚未公布。
Kimi scored an impressive 95.8 on Mazur's Extended NYT Connections, ranking 3rd best, but it cost more than Fable to run. As of last check, its other scores are not in yet.
Kimi 在谄媚测试中得分几乎达到 Claude 水平,'你说得完全正确!'
Kimi scores almost Claude-level on the sycophancy test, 'You're absolutely right!'
我想知道这在多大程度上归因于知识蒸馏:
I wonder how much of this is because of the distillations:


Kimi K3 在现实世界中的表现如何?能否经受住考验并转化为实际应用?
How well does Kimi K3 hold up and translate into the real world?
Kimi K3 到底有多好?我们如何将其置于上下文中?
Exactly how good is Kimi K3? How do we put this in context?
你绝对不能仅仅根据基准测试就说某物是,例如,一个“无可争议的前沿模型”。
You absolutely cannot say something is, for example, an ‘undisputed frontier model’ based on benchmarks alone.
大概存在针对中国共产党不希望你说出的内容的防护措施。
There are presumably safeguards against things the CCP does not want you to say.
我们没有看到防止误用的明确防护措施。在生物学方面是零。
We do not see explicit safeguards to prevent misuse. Zero for biology.
我们既没有看到“看看 Kimi 在生物学上能做什么”,也没有看到“看看 Kimi 在生物学上会拒绝为我做什么”。
We see neither 'look at what Kimi can do in biology' nor 'look at what Kimi refuses to do for me in biology.'
这要么意味着智能分布参差不齐,要么意味着对生物能力的削弱极其新颖。我倾向于认为其中并不涉及高度创新的削弱,因为如果有的话,他们大概会大肆宣扬。
This implies either jagged intelligence or highly innovative nerfing of bio capabilities. I'm going to assume there are not highly innovative nerfs involved, since presumably they would be bragging about that.
此外,依赖此类防护措施是愚蠢的,因为他们很快会开放模型权重,届时世上最恶劣的人会将这些防护移除。
Also, it would be foolish to depend on such safeguards, since they're going to open up the weights soon, at which point the worst person in the world would remove them.
然后是网络安全方面,Sirius 报告称其在此方面出奇地薄弱。但要在网络安全上表现糟糕,同时在通用编程上表现出色并不容易,因为这两者大体上是同一项技能。而且再说一次,我们没有看到有人触碰到明确的保障措施。
Then there’s cyber, where Sirius reports it is strangely weak. But it’s not easy to be bad at cyber while being good at general coding, since they are largely the same skill. And again, we don’t see people hitting explicit safeguards.
Eric 的假设很有趣,编辑期间 Fable 也着重指出了这一点。如果 Kimi K3 在很大程度上是通过从 Fable 进行蒸馏来训练的,但 Fable 拒绝网络安全类任务,那么这就能解释其在网络安全方面的相对能力不足;不过很多技能仍然会迁移过来,因为它们与常规编程是相同的技能。
Eric’s hypothesis is interesting and was highlighted by Fable during editing. If Kimi K3 is being trained largely by distillation from Fable, but Fable refuses cyber tasks, then that would explain a relative capability deficit on cyber, but many skills would still transfer over since they’re the same skills as regular coding.
许多人对这个特别的技巧印象深刻,但我让 Sol 检查了一下,它并不以为然,甚至猜测 K2.6、Gemini 和 GLM-5.2 很有可能在这里也能做到同样的事情。
Many were impressed by this particular trick, but I had Sol check it out, and it was not so impressed, including guessing that K2.6, Gemini, and GLM-5.2 had a good shot at matching its work here.
我之前已经从网络风险的角度解释过,为什么 Mythos 与 Sol 属于不同的类别。关键不在于能够针对某个具体任务做到什么,而在于能够把所有事情整合起来,并大规模自主地执行。Kimi K3 或许属于 Sol 的类别,也或许不属于,现在下结论还为时过早。但显然,它不属于 Mythos 的类别。
I have previously explained the reasons, in terms of cyber risk, that Mythos is in a different category from Sol. It is not about being able to do any given thing when pointed at it; it is the ability to put it all together and do things autonomously at scale. Kimi K3 may or may not be in Sol’s category. Too soon to be sure. It clearly is not in that of Mythos.
许多人非常执着于不去理解这一点,也不去理解一个无法承受假阴性的系统最终会产生一些假阳性。例如,David Sacks 在这里引用 calle 和 clem 的话,称 Kimi K3 完成了一项 Sol 和 Fable 拒绝执行的防御性任务,并由此得出结论:我们应该让我们的模型愿意去做 Kimi K3 能做的任何任务。
Many people are very dedicated to not understanding this, and also to not understanding that a system that cannot afford false negatives will end up with some false positives, such as David Sacks here quoting calle and clem saying Kimi K3 did a defensive task Sol and Fable refused to do, and thus concluding that we should just have our models be willing to do any task the Kimi K3 can do.
我的意思是,是的,如果我们能有这样的护栏——只阻止坏任务、帮助好任务,或者只拒绝那些既可能有害又无法以其他方式完成的任务——那当然很好。但事实证明这很难。Anthropic 绝对应该改进他们的分类器和护栏,以减少假阳性;但 OpenAI 的分类器已经相当合理了,而且是的,这确实会带来一些假阳性,尤其是当你引用那些最不擅长应对假阳性的人时。
I mean, yes, it would be great if we could have guardrails that stopped only the bad tasks and helped with the good tasks, or only refused the tasks that were both plausibly bad and that also could not otherwise be done. But it turns out that is hard. Anthropic should absolutely improve their classifiers and guardrails to reduce false positives, but OpenAI’s classifiers are pretty reasonable, and yes that will involve some false positives, especially if you quote those who are least able to work around it.
正如发布周末常见的情况,服务器似乎过载了。这个模型并不容易提供服务,而且算力有限。Moonshot 的应对措施是暂停新订阅以优先保障现有成员,并正在努力增加容量。
As is often the case on a release weekend, servers seemed overloaded. This is not the easiest model to serve and compute was limited. Moonshot is responding by pausing new subscriptions to prioritize current members and is working to add capacity.
一旦该模型可由其他方提供服务,供应预计将更好地匹配需求。
Supply will presumably better match demand once the model can be served by others.
与此同时,要获得响应并不容易。
In the meantime, getting a response has not been easy.
这个模型也并非那么便宜,尽管他们显然收费不足。
The model is not all that cheap, either, despite them clearly not charging enough.
对于以高品质方式呈现的高品质产品,你可以收取可观的溢价,但没有人拥有允许数量级加价的空间,即使在纯粹的双头垄断中也不会。
You can charge a solid markup for a high quality product presented in a high quality way, but no one has a gap that allows orders of magnitude of markup, and they wouldn’t even in a pure duopoly.
最好的模型仍然可能价值不菲。对于很大一部分 token 而言,如果只能在这些选项中选择,我宁愿为 Sol 或 Fable 支付相当于 GPT-5.5 或 Opus 基础成本十倍的价格。即使 Kimi 与 GPT-5.5 水平相当,它也会落败。
The best model can still be worth quite a lot. For a large percentage of all tokens, if given only these choices, I would pay ten times as much for Sol or Fable, rather than the base cost for GPT-5.5 or Opus. Even if Kimi is on par with GPT-5.5, it loses out.
这种定价意味着,你是在将 Kimi 订阅与 Claude 或 ChatGPT 订阅进行比较——后者的价格都在 20 至 200 美元之间——而 Kimi 的限额似乎并不那么高。
The pricing means that you're comparing Kimi subscriptions to Claude or ChatGPT subscriptions, which all go from $20 to $200, and Kimi's limits don't seem that high.
(本节完全写于今天关于特朗普政府的新闻发布之前。)
(This section was entirely written prior to today's news regarding the Trump admin.)
K3 不是 Mythos。这并不意味着 Kimi K3 是安全的开放权重发布。
K3 is no Mythos. That does not mean that Kimi K3 is a safe open weights release.
Kimi K3 有望成为最强大的开放权重模型。其他模型可能在特定任务上效率更高,但在我们最担心的网络能力上,以及大概也在我们最担忧的生物能力上,K3 很可能是迄今最强的开放模型。我们应当有多担心?
Kimi K3 is poised to be the most capable open weights model. Others might be more efficient for a task, but on the cyber capabilities we worry about most, and presumably also on the bio ones we'd worry about most, K3 is probably the strongest open model yet. How worried should we be?
我相信我们应当有非零的担忧,认为会出现实质性的麻烦。中等结果是我们会看到某些形式的“普通可应对”麻烦略有上升,不太有趣但也不会超出我们的处理能力,而且事后看也不会让我们希望当时停止发布。但这里存在尾部风险。
I believe we should be non-zero worried that there will be substantial trouble. The median outcome is that we see modest upticks in some forms of 'ordinary decent' trouble, not fun exactly but nothing we cannot handle, and nothing that would in hindsight make us want to have halted release. But there is a tail risk here.
我估计大约有 10%的可能性我们会后悔让这件事发生,还有大约 2%的可能性这是一个相当严重的错误。
I'd estimate something like a 10% chance we regret letting this happen, and ~2% chance that it was a rather serious mistake.
它几乎肯定不会造成“广泛的社会混乱”。
What it will almost certainly not do is cause 'widespread societal chaos.'
这与发布 Mythos 或与 Mythos 相当的模型形成对比。这一差距非常重要,我已多次尽力解释为什么 Mythos 在此具有独特性,希望与 Mythos、Fable 和 Sol 共处的时间能帮助我们做好准备。
This is in contrast to releasing Mythos, or a model on par with Mythos. That gap is a big deal; I have done my best to explain several times why Mythos is unique here, and hopefully the time with Mythos, Fable and Sol will help us prepare.
1. 认识到前沿 AI 模型存在严重的安全隐患。
1. Recognize that there are serious security concerns with frontier AI models.
2. 知道他们在前沿能力和算力上落后,并认识到开放性带来的优势,无论是在传播还是声望积累方面。
2. Know they are behind on frontier capability and compute, and recognize the advantages they get from openness, both in diffusion and aura farming.
3. 追求有意识的快速跟随策略,专注于效率和传播,以美国实验室为先导,并经常使用蒸馏。这仍然涉及创新,且往往意味着在某些方面做得更好。转向“领先”将是一个巨大、困难且昂贵的转变,即使美国在相关意义上完全“暂停”,也需要一段时间。
3. Pursue an intentional fast following strategy, focused on efficiency and diffusion, using American labs to lead the way and often using distillation. That still involves innovations, and often means doing some things better. Transitioning to 'taking the lead' would be a huge, difficult and expensive transition, which would take a while even if America fully 'paused' in the relevant senses.
4. 并未直接推动阿里巴巴开源 Qwen。他们继续支持开放性,而阿里巴巴因选择封闭而失败,因为其模型不足以在封闭市场中竞争。所以阿里巴巴退缩了。
4. Did not directly push Alibaba to open up Qwen. They continue to support openness, and Alibaba’s experiment with being closed failed because their models are not good enough to compete for the closed market. So Alibaba folded.
1. 我没有内部消息,也不完全确定,但这是我的猜测。
1. I don’t have insider info and I’m not certain, but this is how I’d bet.
5. 尚未完全被 AGI 洗脑,也没有完全经历他们的“神话时刻”。
5. Are not yet so AGI pilled and have not fully had their Mythos Moment.
6. 他们计划尽可能长久地搭乘开放性浪潮,但做好准备,就像在新冠疫情期间那样,一旦不得不收手,就会迅速且严厉地收手。
6. Plan to ride the openness wave as long as they can, but are prepared, as they did in COVID, to come down and come down hard the moment they have to.
7. 他们主要通过事先限制和对中国模型的其他规定来确保准备就绪,这种监管和控制程度,若是那些“捍卫开源”的人真正了解正在发生的事情,一定会大喊大叫;而你也建议在美国实施类似的规定。
7. Ensure their preparedness largely via prior restraints and other rules on Chinese models, with a level of regulation and control that would have most who are ‘defending open source’ screaming bloody murder if they actually understood what was going on, and you suggested applying similar rules in America.
我认为,鉴于他们的立场,这是一个非常合理的策略。如果我们还把他们当前对 AGI 的信奉程度视为既成事实,那么这对他们来说显然是正确的做法。
I think this is a highly reasonable strategy, given their position. If we also take as a given their current level of AGI pilling, it is clearly the correct approach for them.
Nathan Lambert 五月份在中国实验室内部发表了一些看法。我不太愿意认同文化概括,但这个故事似乎经得起推敲。
Nathan Lambert offered thoughts back in May from inside China’s labs. I hesitate to endorse cultural generalizations, but the story seems like it checks out.
Dean Ball 发表了一条很好的评论,似乎值得完整分享,这不仅因为他在白宫的经历和他在 OpenAI 的新职位,部分是因为它很有思想,部分是因为那些回应完全离谱。
Dean Ball had a good comment that seems worth sharing in full, including because of his history at the White House and his new position at OpenAI, in part because it is thoughtful, and in part because the responses are absolutely unhinged.
随后,就在截稿之际,Dean Ball 被证明是对的,正如下一节所述。除这一段外,我没有对本节进行任何编辑,只是加入了与 Emil Michael 的交流。这不是事后诸葛。
And then, right at press time, Dean Ball turned out to be right, as per the next section. Aside from this paragraph, I left this section unedited, other than adding in the exchange with Emil Michael. It was not written with hindsight.
如果你认识 Dean Ball,你就会知道这正是他在加入 OpenAI 之前在类似情况下会发表的那种评论。这与他在公开和私下之前的想法完全一致。
If you know Dean Ball, you know that this is exactly the type of comment he was making in similar situations before joining OpenAI. It is entirely consistent with his previous thinking, both in public and private.
习近平昨天刚发表讲话,支持开放模型,但也强调了控制的需要,而 Kimi K3 正在接近矛盾变得明显的临界点。
Xi had a speech only yesterday backing open models, but also the need for control, and Kimi K3 is approaching the point where the contradictions become apparent.
我同意中国共产党尚未理解这一局面,对 AGI 的接受程度也不够,但我也认为,站在他们的立场上,如果我对 Kimi K3 最终落点的判断是对的,这是一个我原本预期他们会承担的经过计算的风险。
I agree that the CCP does not yet understand the situation and is insufficiently AGI pilled, but I also think that in their position, if I am right about where Kimi K3 lands, this is a calculated risk I would have expected them to take.
下一轮可能就是他们需要做出更艰难决定的时候。
The next round is where they may have to make a more difficult decision.
Dean Ball 讨论的加速或减速,指的是最大前沿模型的能力,而不是能力的扩散或芯片的使用。很多人没有理解这一点。
Dean Ball is discussing acceleration or deceleration as being about the capabilities of the largest frontier models, not about the diffusion of capabilities or use of chips. A lot of people did not understand this.
许多‘加速主义者’并没有连贯的世界模型,当然也不理解二阶效应,他们大多只是在感受那种更开放模型、无限制和不可治理性的加速氛围。而且,因为他们认同开放和不可治理性,便把这些与他们认为好的事物(包括加速)联系起来。
A lot of the ‘accelerationists’ do not have coherent world models and certainly do not understand second order effects, and are mostly vibing the acceleration of more open models and no restrictions and ungovernability. And because they vibe openness and ungovernability, they associate it with things they think are good, which include acceleration.
此外,他们真正想要‘加速’的,往往只是自己的公司、产品和玩具,而不是整个 AI。他们并不认真对待 AGI,多半只想做出酷炫的东西。因此,他们想要酷炫的玩具来帮助构建自己的酷炫事物,并在此之上进行开发,或者希望这些玩具能‘加速’销售。这非常能引起共鸣。
Also, what they actually want to ‘accelerate’ for real is often their own companies and products and toys, not AI in general. They don’t take AGI seriously and mostly want to build cool things. So they want cool toys to help build their cool things, and to build on top of those toys, or they want the toys to ‘accelerate’ sales. Highly relatable.
显然,开放模型在短期内是能力扩散的加速主义者,这总体上是好事。
Clearly open models are short term accelerationist for diffusion, which is net good.
但确实,开放模型有一个直白的理由就是加速前沿,因为它们让每个人都能在既有成果上继续构建。这当然有助于其他人赶上前沿。我仍然认为这在历史上很重要。你的开放发布加速了其他人所拥有的东西。
But also, yeah, there’s a straightforward case for open models being acceleration of the frontier, as they let everyone build on everything. Certainly it helps others catch up to the frontier. I continue to believe this was important historically. Your open release accelerates what others have.
从某种意义上说,这是减速主义的,因为它降低了创新带来的财务回报。我同意,如果你考虑例如 Plan A 式的前沿模型强制开放,这种效应很可能是主导性的。
It is decelerationist in the sense that it reduces the financial benefits to innovation. I agree that this is likely the dominant effect if you considered, e.g., a Plan A style mandatory openness of frontier models.
这进一步证明,那些‘加速主义者’中有很多人——而且是最突出的一些人——是伪善的:他们想要监管俘获、公共资金和能让自己获胜的规则;他们对别人的指责有一部分是投射,因为那正是他们自己会做的事。他们会痛斥给其他所有人——包括竞争对手和有需要的人——的施舍和政府帮助,还威胁要‘带着球离场’,仿佛自己是安·兰德小说里的英雄。但随后他们却向政府索要施舍。
This is more evidence that many—and many of the most prominent—of those ‘accelerationists’ are hypocrites: they want regulatory capture, public funds, and rules that make them win; and their accusations against others are in part projection, because it is what they would do. They will rail against handouts and government help for everyone else, both their competitors and for people in need, and also threaten to take their ball and leave, like heroes in an Ayn Rand novel—except then they ask for the government handouts.
他们不认为提供前沿模型是一项‘合法业务’,因为那不是他们的业务,他们没在其中投资,因此它就是不合法的。就这么简单。这或许有助于解释 Marc Andreessen 那个出了名的奇怪错觉:他直到今天还声称,Biden 政府当面告诉 Marc Andreessen,说它‘不会允许存在 AI 初创公司’。
They don’t think of serving a frontier model as a ‘legitimate business’ because it is not their business; they are not invested in it, ergo it is illegitimate. Simple. This perhaps helps explain Marc Andreessen’s famously bizarre delusion: he to this day claims the Biden administration told Marc Andreessen, to his face, that it would ‘not allow there to be AI startups.’
同样,这里是 Will Manidis 把 Ball 的帖子解读为呼吁美国‘清除美国市场上更便宜的前沿竞争者’——并且因此走红。抱歉,什么?然后是大卫·萨克斯(David Sacks),即使按他的标准也异常地不诚实,假装不明白许多我愿意相信他其实理解的事情。
Similarly, here is Will Manidis interpreting Ball’s post as calling for America to ‘clear the American market of a cheaper frontier competitor’—and going viral for it. Sorry, what? And here is David Sacks being unusually disingenuous even for David Sacks, pretending not to understand many things I like to think he understands.
我们甚至把大家最喜欢的那位有倾向性的战争部副部长也拉了进来——他曾帮助将 Anthropic 宣布为供应链风险。看来这是为了支持他的立场,即在一份政府合同中,使用 Kimi K3 应该比使用 Claude 更容易。
We even got everyone’s favorite tilting Undersecretary of War, who helped declare Anthropic a supply chain risk, in on the act. It would seem this is to back up his position that it should be easier to use Kimi K3 in a government contract than Claude.
另据报道,“我永远不会离开这个应用”,而且我们都需要好好笑一笑:
In other news, “I Am Never Leaving This App,” and we all need a good laugh:
整个社区——也就是所谓的“反 1047 联盟”——作为开发者和风投社区中的特定子集,表明它会将任何关于“政府可能不鼓励在美国关键基础设施或关键供应链中使用中国开源模型”的建议视为越界,即使那仅仅是对已经明确正在发生的事情的预测。它还表明,这将被视为一种“监管俘获”的企图。
An entire community—what one might call the ‘anti-1047 coalition’—a certain subset of the developer and VC communities, revealed that it would treat as beyond the pale any suggestion that the government might discourage use of Chinese open models in American critical infrastructure or our key supply chains, even if that suggestion was merely a prediction of what is already clearly in the process of happening. It also revealed that this would be treated as an attempt at ‘regulatory capture.’
这被一些人称为一个澄清时刻,尤其是对于那些此前认为这类人理性、务实、爱国且具备阅读理解能力的人来说。他们对“什么能力会改变你的想法”的回答是:没有。
It was what some would call a clarifying moment, especially for those who previously thought such people were reasonable, practical, patriotic, and had reading comprehension. Their answer to ‘what capabilities would change your mind’ is none.
到了某个时候,网络蜂群终将被认清本来面目。
At some point the online swarm is recognized for what it is.
他们就是我们原本以为的那种人。而到目前为止,我们一直让他们免于责罚,试图安抚他们,并让他们在很大程度上主导了美国的 AI 政策。
They are who we thought they were. And, until now, we let them off the hook, tried to placate them, and let them drive a good deal of American AI policy.
有理由反驳说,开放软件往往是由大公司作为基础设施提供的,比如谷歌对安卓(Android)的做法。我的推测是,以当前的投资和资本支出水平,这笔账根本算不过来。
It is reasonable to push back that often open software is provided by major corporations as infrastructure, such as Google does with Android. My presumption is that the math would not be mathing at current levels of investment and capex spending.
我甚至不确定他们是否需要采取任何行动。风险是现实存在的。如果我是一家与政府或关键基础设施打交道的严肃行业中的老牌企业,我可不会因为别人知道你用的是中国模型而感到兴奋。
I’m not even sure they have to do anything at all. The risks are present. If I was an established corporation in a Serious Business that dealt with the government or critical infrastructure and such, I would not be excited by the problems of others knowing you were using Chinese models.
可惜,人们因为这条和其他帖子对迪恩·鲍尔进行了猛烈抨击,还把他的推文当作代表 OpenAI 的官方沟通策略,满是“你当然会这么说,因为<OpenAI>”之类的论调。这意味着今后他不能再以同样的方式与我们分享这类帖子了,因为不仅那样做毫无乐趣,还会破坏反馈循环,而这正是发布此类推文的意义所在。
Alas, people came at Dean Ball hard for this and other posts, and also acted as if his Tweets were official communications strategy on behalf of OpenAI, with lots of ‘of course you said that because <OpenAI>.’ Which means he won’t be able to share such posts with us in the same way going forward, because not only is that absolutely no fun, it also invalidates the feedback loops that were part of the whole point of Tweeting such things.
迪恩还提供了一份更具体的复盘,说明那条帖子究竟出了什么问题,鉴于他的新职位,他具体不能再做什么。随后他还阐述了自己对开源软件的立场,并首先解释这源于对开放性的深沉热爱:
Dean also offered a more specific post-mortem on what went wrong with that particular post, in light of his new position, and what exactly he can no longer do. He also then lays out his position on open source, after explaining this comes from a place of deep love for openness:
如果仅仅是他回复中的信号难以找到,我会建议 Dean 坚持下去,但这种程度的敌意需要付出更高的代价。我认为在 Twitter 上硬撑下去是不可持续的。希望他能调整。
If it was merely that the signal in his replies was hard to find, I would advise Dean to power through, but this level of hostility comes with a higher price. I do not think pushing through it is sustainable on Twitter. Hopefully he can adjust.
我非常幸运,我所面临的敌意是极其有限且可控的。
I am very blessed that I have faced a highly modest and manageable amount of hostility.
在大型实验室工作的人,以及大多数其他人,总是会有一定的偏见。你在哪里工作会影响你的思维方式和你选择说的话。你必须对此进行调整。但我能够将 Barak、Achaim 以及 OpenAI 的许多其他人,尤其是 Roon 和现在的 Dean Ball,视为主要是在表达他们真实的想法,并且只有在他们明确表示代表 OpenAI 发言时才代表 OpenAI。
There will always be some amount of bias from those who work at a major lab, and from most other people as well. Where you work colors how you think and what you choose to say. You do have to adjust for that. But I have been able to treat Barak, Achaim and many others at OpenAI, and especially Roon and now Dean Ball, as primarily saying what they actually think, and only speaking on behalf of OpenAI when they explicitly say they are doing so.
我绝对不赞成这样做,Dean Ball 也不赞成,但现实就是如此。
I am absolutely not in favor of this, and neither is Dean Ball, but here we are.
预测白宫会做什么,或描述它正在考虑做什么,与你认为我们应该做什么是非常不同的。
A prediction of what the White House will do, or a description of what it is considering doing, is very different from what you think we should do.
我确实认为,我们至少应该考虑将中国的开放模型视为供应链风险,并采取诸如将其排除在关键基础设施之外等措施,但这与看起来正在考虑的事情不同。
I do think that we should at least consider treating Chinese open models as supply chain risks, and doing things like keeping them out of critical infrastructure, but that is different from what it looks like is being considered.
实际上想要‘禁止开放模型’的从来都不是你想象的那些人。请记住,所有出于 AI 安全动机的法案都谨慎地尽量减少对开放模型的影响,而白宫的行动将试图最大化影响。这是不同的世界。
The ones who actually want to ‘ban open models’ are never the ones you think. Remember that all the AI-safety-motivated bills were careful to minimize impact to open models, whereas the White House move will be attempting to maximize impact. Different worlds.
我早就开始认为‘谷歌不再处于第一梯队’了,但有很多健康的竞争并非来自中国。Curi 直接从 David Sacks 对 Dean Ball 原始推文的虚伪误读中拿来了‘双头垄断’和‘锁定’的说法。OpenAI 和 Anthropic 并不是推动这一点的力量。
I was early on the ‘Google is no longer in the top tier’ train, but there is plenty of healthy competition that is not Chinese. Curi is taking the ‘duopoly’ and ‘lock in’ lines directly from David Sacks’s disingenuous misreading of Dean Ball’s original tweet. OpenAI and Anthropic are not driving this.
白宫跟往常一样,提议以一种极其生硬的方式来处理这件事。
The White House, as per usual, is proposing to do this in a maximally blunt way.
Dean Ball 的预测——同时也是在暗示如果白宫认为有必要这么做,应当如何操作——就是简单地制造监管不确定性,这足以阻止大型企业在关键环节使用外国模型,同时让初创公司和开发者们自行其是。再往上一层便是将其标记为“供应链风险”。
Dean Ball’s prediction, which was also a subtle hint as to how to do it if the White House decided it needed to do it, was simply to create regulatory uncertainty, which would be sufficient to discourage big players from using foreign models in critical places, while letting startups and builders have their fun. A ‘supply chain risk’ designation would be the next step up from that.
这项提案则是另一回事。这简直是一记大锤。
This proposal is something else. This is a sledgehammer.
这是我私下听说过的过去的一项提案,实际上等于禁止云服务商提供这些模型。这也极其愚蠢,只会把业务推向竞争对手,却一事无成。
This is a past proposal I’d heard about privately, and would be a de facto ban on the cloud providers serving those models. This would also be deeply stupid, driving business to the competition without accomplishing anything.
我确实理解 David Sacks 不得不不断反击此类过度反应的处境。这并不能为他的行为开脱,但几乎每个政界人士都有自己的麻烦和更疯狂的人要应对。
I do sympathize with David Sacks that he had to keep pushing back on overreactions like this. That doesn’t excuse his actions, but almost everyone in politics has troubles and crazier people of their own to deal with.
这正是 Dean Ball 的预测。这个选项现在看起来是不是开始变得相当不错了,对吧?
That’s exactly the Dean Ball prediction. That option might be starting to look pretty good right around now, huh?
好的。回到 Kimi K3 的实际能力。
Okay. Back to Kimi K3's actual capabilities.
那将是这里潜在情景的上限。
That would be the top end of potential scenarios here.
我强烈同意 Roon 的第二段。第一段至少有点操之过急;另外,请注意他只谈论‘公开’模型。
I strongly agree with Roon's second paragraph. The first one at least toys with jumping the gun; also, notice he is only talking about 'public' models.
正如上文 Dean Ball 所说,它对大多数人的智能体式编码显然非常好用,可能与 2026 年第一季度的模型相当。大多数智能体式编码与基准测试和训练任务所衡量的内容相当接近,所以你可以相对“浅层”,却仍然在日常工作中表现出色。
As per Dean Ball above, it is clearly very good for most people's agentic coding, plausibly on par with models from Q1 2026. Most agentic coding is rather close to what benchmarks and training tasks measure, so you can be relatively 'shallow' and still impress in the day to day.
Tushit 运行了一个内部 React/前端评测,发现 Kimi 与 Opus、Sonnet 和 Grok 4.5(?)相比是最慢的,但成本大约是 Opus 的一半,而且它们通常都能成功,实际上 Grok 更胜一筹。这听起来像一个已饱和的基准测试,但有些人的真实世界任务也已饱和。处理好普通事务也很重要。
Tushit runs an internal React/frontend eval, finds Kimi the slowest versus Opus, Sonnet and Grok 4.5 (?), but about half the cost of Opus, and all of them usually succeed, with Grok actually coming out ahead. Sounds like a saturated benchmark, but some people's real world tasks are saturated. Handling the ordinary stuff matters, too.
因此,有些人坚持使用这些已饱和的基准测试:
Thus, some people stick with the saturated benchmarks:
此外,我们还看到了一些看起来很酷的 3D 内容。
Also we've seen a bunch of 3D stuff that looks cool.
不过,下面是相反的意见——反应总是各不相同:
Here is the opposite opinion, though, reactions always vary:
最重大的说法是,它能兑现其基准测试所显示的能力。Elanor 明确提出了这一主张,尽管几乎所有人都不同意。
The biggest claim would be that it lives up to its benchmarks. Elanor is explicitly claiming this, although almost everyone else disagrees.
无意的。对。让我们看看 J 空间。
Inadvertent. Right. Let’s see the J-space.
这是一个好模型的标志。但这并不意味着有很强的理由选择 Kimi K3 去做那件事。你不会得到那么大的折扣。
This is the sign of a good model. That doesn’t mean there is any strong reason to choose Kimi K3 to do that thing. You are not getting that large a discount.
这在我看来基本正确,但 Kimi K3 可能已经足够好,因此在某些领域(例如,如果 Harvey 的结果成立)它会处于顶尖位置并被选中:
This seems mostly right to me, but Kimi K3 is probably good enough that there will be some areas (e.g. if the Harvey result holds) where it is at the top and gets the call:
Kimi K3 声称自己是 Claude 的频率有多高?不常见,但有时会。
How often does Kimi K3 claim to be Claude? Not usually, but sometimes.
[](https://substackcdn.com/image/fetch/$s_!672p!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F507334e3-b561-4976-a193-17a395309611_1200x600.png)
[](https://substackcdn.com/image/fetch/$s_!672p!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F507334e3-b561-4976-a193-17a395309611_1200x600.png)
除了“声称自己是 Claude”的问题之外……
In addition to the ‘claims to be Claude’ issue…
这里的“它”指的是超常规的基准测试性能提升。
By “it” we mean outsized benchmark gains.
节省 GPU 算力使我们能够扩大模型规模,进而带来能力提升,所以是的。
GPU savings enable being larger, which enables capability gains, so yes.
我考虑过 Teortaxes 的假说,但已予以否定。中国无疑也拥有大量杰出人才,他们完全能够做出创新,尤其是在效率方面。但如果你的理论要求中国人在总体上比美国人更擅长 AI 研究,尤其是考虑到可用资源和经验,我认为这并不可信。
I have considered and rejected Teortaxes’s hypothesis. China doubtless also has lots of great talent. They absolutely can do innovative things, especially in terms of efficiency. But if your theory requires the Chinese to be better AI researchers, in general, than the Americans are, especially if this takes into account available resources and experience, I don’t think that is credible.
总体而言,我推测中国模型大量涉及“benchmaxxing”(刷基准)、“usemaxxing”(刷使用),浅层泛化,关注相对优势,以及对 Claude(或 GPT,但在这种情况下我们可以确信是 Claude)的蒸馏。
In general I presume Chinese models involve a lot of benchmaxxing, usemaxxing, shallow generalization, focus on relative strengths, and distillation from Claude (or GPT, but in this case we can be confident it was Claude).
奖励黑客与作弊是现实存在的可能性,但目前难以评估。
Reward hacking and cheating are live possibilities but hard to assess for now.
Kimi K3 也从扩大规模中受益。
Kimi K3 also benefits from moving up in size.
Kimi K3 可能是截至目前中国发布的最令人印象深刻的模型,单就纯能力而言。它是一个非常好的模型。我目前的猜测是,Kimi K3 的实际表现会比其亮眼的基准测试成绩略逊一筹,但在某些相对高性能的领域它具有竞争力,并且有一种独特风格,有些人会喜欢。它距离 Fable 还很远,我也不认为它距离 Sol 有那么近。
Kimi K3 is potentially the most impressive Chinese release so far in terms of pure capability. It is a very good model. My current guess is that Kimi K3 will modestly underperform its highly impressive benchmarks, but with some areas of relatively high performance where it is competitive, and with a unique style some people will enjoy. It is not close to Fable, and I do not believe it is that close to Sol.
如果今天发布其权重,它将是能力最强的开放模型。他们或许应该抓紧,因为新的 Qwen 即将推出,预览版已经上线(但我还没有看到任何人尝试过它的报告),它很可能比 K3 更好,所以我们很可能很快又要再来一遍。
If its weights were released today, it would be the most capable open model. They might want to hurry, since a new Qwen is dropping soon, with the preview live (but I have seen zero reports from anyone trying it) which might well be better than K3, so chances are we will soon all have to do this over again.
在权重发布并且我们有更多时间之前,我们不知道它有多好。目前访问时好时坏且受限,我们还有很多不了解的地方。
We do not know how good until the weights are released and we have more time. For now access has been spotty and limited, and there is much we do not know.
和往常一样,有些人变得忘乎所以,他们说最新发布改变了一切,美国的领先地位已经消失,中国实验室现在‘赢了’,开放模型将会‘赢’,现在对美国模型的所有限制都很愚蠢,等等。不要成为那样的人。
As usual, there are some who are getting carried away, who say the latest release changes everything, that the American lead is gone, that Chinese labs are now ‘winning,’ that open models will ‘win,’ that all limits on American models are foolish now, and so on. Do not be one of those people.
你也不应回避发布能力越来越强的开放权重模型所带来的安全和其他安全方面的担忧。现在还不是‘Mythos 级开放模型’的时刻。我预计这次最多只会造成适度的冲击,而且这些冲击会逐渐发生,并伴随一些较小的尾部风险。
Nor should you shrink from the security and other safety concerns of releasing increasingly capable open weights models. This is not the ‘Mythos-level open model’ moment. I expect at most modest disruptions this time, and for those to occur gradually, with some small tail risk.
但确实,除非中共干预加以阻止,否则我们应预期到年底会有一个模型跨越这一阈值。人们不能简单忽视其中涉及的风险,美国政府也不能,他们同样不能忽视这些模型是中国的这一事实。行动和限制即将到来。那些无法接受这一点、甚至无法接受有人指出这一事实,同时又炒作 Kimi K3 并为之叫好的人,正在制造一个澄清的时刻。
But yes, absent CCP intervention to stop it, we should expect a model to cross that threshold by the end of the year. One cannot simply ignore the risks involved in that, and the American government cannot either, nor can they ignore the fact that these models are Chinese. Actions and restrictions are coming. Those who cannot accept this, or even accept people pointing this fact out, while simultaneously hyping up Kimi K3 and cheering it on, are creating a clarifying moment.