Kimi K3:开放权重的升级

Kimi K3: The open-weights escalation

内森·兰伯特 Nathan Lambert · Allen Institute for AI · 2026-07-20 · Interconnects ↗

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

本文分析了月之暗面公司发布的 Kimi K3——一个拥有 2.8 万亿参数的开源权重 MoE 模型,及其对 AI 生态系统的影响。文章认为,K3 显著缩小了开源与闭源模型之间、以及中美实验室之间的性能差距,将其从 6-9 个月缩短至 3-5 个月。核心论点是,以月之暗面为代表的中国 AI 实验室不仅仅是依赖蒸馏的快速追随者,而是能够进行前沿创新,这得益于强大的文化和资本效率。文章还讨论了中国对开源 AI 的战略承诺,这体现在习近平在 WAIC 上的主旨演讲中,以及开源模型对闭源实验室的经济减速效应。结论是,虽然开源模型减缓了对前沿实验室的投资,但它们加速了 AI 在经济中的扩散,带来了社会效益,并为解决安全挑战赢得了更多时间,尽管美国仍有望在前沿能力上保持领先。

This article analyzes the release of Moonshot AI's Kimi K3, a 2.8T parameter open-weights MoE model, and its implications for the AI ecosystem. It argues that K3 significantly narrows the performance gap between open and closed models, as well as between Chinese and American labs, reducing it from 6-9 months to 3-5 months. The core thesis is that Chinese AI labs, exemplified by Moonshot, are not merely fast followers relying on distillation but are capable of frontier-level innovation, driven by strong culture and capital efficiency. The article also discusses China's strategic commitment to open-source AI, as signaled by Xi Jinping's WAIC keynote, and the economic decelerationist effect of open models on closed labs. It concludes that while open models slow investment in frontier labs, they accelerate AI diffusion across the economy, offering societal benefits and more time to address safety challenges, though the U.S. is still expected to lead in frontier capabilities.

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 7)

全文 · Full text(逐段中英对照)

对 AI 生态系统的全球影响 The global implications on the AI ecosystem.

7 月 16 日(星期四),Moonshot AI 发布了其最新旗舰模型 Kimi K3。K3 是一个 2.8T 参数的 MoE 模型,其权重将于 7 月 27 日发布。本文的大部分内容是在假设 Moonshot 信守权重发布日期承诺的前提下,对生态系统状态的反思。这是一种对均衡状态的更极端看法;如果实际情况是中国拥有同样强大但不开放的模型(即 K3 从未发布),那么许多结果最终会落在中间地带。

On Thursday, July 16th, Moonshot AI released their latest flagship model Kimi K3. K3 is a 2.8T parameter MoE model which will have its weights released on July 27th. Much of this article follows as a reflection on the state of the ecosystem, under the assumption that Moonshot keeps their promise of the weights release date. This is a more extreme view of the equilibrium, and many of the results end up in a middle ground if the state of affairs is that China has similarly powerful, but closed models (i.e. K3 is never released).

关键事实是,无论是开源到闭源之间的差距,还是美国到中国之间的模型性能差距,都已从争论中的 6-9 个月缩短到更短的时间,比如 3-5 个月。

The key fact is that either the open-to-closed or American-to-Chinese model performance gap has been reduced from the debated 6-9 months to something shorter, say 3-5 months.

从发布材料来看,K3 显然是一个真正的前沿模型。它将是自 DeepSeek R1 以来,开源模型最接近前沿的一次。DeepSeek R1 是另一番景象:那是一家中国实验室极其迅速地转向推理模型,并比许多美国公司更快地发布了一个推理模型。Kimi K3 则是一个中国实验室在扩展已知领域(数据、算法、架构、工具、环境等)时执行力的体现。

From the release materials, it is clear that K3 is a true frontier model. It will be the closest open models have been to the frontier since DeepSeek R1. DeepSeek R1 was a different story. This was a Chinese lab being extremely quick to pivot to reasoning models and release one faster than many American companies. Kimi K3 is an example of a Chinese lab executing on scaling the known areas: data, algorithms, architecture, tools, environments, etc.

Kimi K3 在 Vals AI 指数中总体排名第 2,在 Artificial Analysis 的智能指数中总体排名第 3(仅被 Claude Fable 和 GPT-5.6 Sol Max 击败,但价格更低),在 Frontend Code Arena 中总体排名第 1,以及其他更令人印象深刻的成绩。Moonshot AI 正以远远更少的资源与 Anthropic 和 OpenAI 正面交锋。

Kimi K3 comes in at #2 overall on the Vals AI index, #3 overall on Artificial Analysis's Intelligence Index (only beaten by Claude Fable and GPT-5.6 Sol Max while being cheaper), #1 overall in Frontend Code Arena, and more impressive results. Moonshot AI is going toe to toe with Anthropic and OpenAI with far, far fewer resources.

这显然是有史以来最强的开源模型。看清楚这个模型后,应当明白:即便来自美国闭源前沿模型的对抗性蒸馏有所贡献,其作用至多也是相对较小的。那些追随蒸馏恐慌并得出错误结论——认为中国 AI 实验室只是因为窃取知识产权才产出好模型——的 AI 观察者们即将醒悟:中国公司在构建模型方面与美国领先公司同样出色。Moonshot AI 正在解决 OpenAI 或 Anthropic 的人正在解决的许多相同问题。我相信关于蒸馏的讨论和压力还会继续,但证据已经表明,中国公司不仅能做快速追随者。

It is clearly the strongest open model ever released. It should be clear looking at this model that if adversarial distillation from the closed frontier models in the U.S. contributed, it is at most to a relatively small degree. AI observers who followed the distillation panic and came away with the wrong conclusion that Chinese AI labs are only producing good models due to IP theft are in for an awakening – that Chinese companies are extremely good at building models in the same way the leading American companies are. Moonshot AI is solving many of the same problems that folks at OpenAI or Anthropic are solving. I'm confident there will be more distillation discussion, and pressure, but the evidence is now out that Chinese companies can do more than just fast following.

在我前往中国期间会见了 Kimi 核心团队的一些成员,我清楚地感受到他们拥有令人难以置信的文化,有人称之为“气场”,以及在 GPU 受限的环境约束下表达这种文化的自由。既然构建模型在很大程度上是一场规模扩张的游戏,构建优秀模型的能力仍主要取决于个人的执行力、动力和表达。访问过他们之后,这一结果就不那么令人惊讶了。我访问过许多 AI 公司,很少有能让你立刻感受到其文化的。

Meeting some of the core Kimi team on my trip to China, it was clear to me that they had incredible culture, some would say aura, and a freedom to express it – within the constraints of a GPU-limited environment. Where building models is so much of a scaling game, much of the ability to build a good model still comes down to individual execution, motivation, and expression. Having visited them, this result is less surprising. Having visited many AI companies, very few have a culture that you can immediately pick up like this.

与此同时,中国的人工智能应用趋势起步晚于美国。因此,尽管所有中国实验室拥有的算力远少于美国同行,但其中更多的算力确实可以用于训练。当我开玩笑地提到 OpenAI 的普通研究员可以拥有多少算力——比如说几千台 H100 等效机器——Kimi 的研究人员感到震惊。Kimi 的组织架构和构建模型的方法无疑反映了这一点,但如果没有大量专有信息,很难梳理出具体的情况。

At the same time, China’s AI adoption trends started later than those in the U.S. So, while all the Chinese labs have way less compute than their counterparts in the U.S., more of it can certainly go to training. When I joked around about how much compute an average researcher at OpenAI could have – say a few thousand H100 equivalent machines – the researchers at Kimi were shocked. The org chart and approach to building the Kimi models surely reflect this, but it is difficult to tease out what this looks like without substantial proprietary information.

顶尖模型性能的状况大致如下:

The state of affairs on peak model performance is roughly as follows:

3. Moonshot AI – Kimi K3(开放权重*)

3. Moonshot AI – Kimi K3 (open weights*)

5. 智谱(Z.ai)– GLM 5.2(开放权重)

5. Zhipu (Z.ai) – GLM 5.2 (open weights)

8. 阿里巴巴 – Qwen 3.7 Max(3.8 已发布,写作时也将开放权重)

8. Alibaba – Qwen 3.7 Max (3.8 announced, also to be open-weights, when writing)

令人惊讶的是,DeepMind 和其他一些美国巨头排名如此之低。从很多方面来看,xAI 团队值得更多赞誉。以下是 Artificial Analysis 的视觉摘要:

It is astonishing to see DeepMind and some of the other American giants this low. In many ways, the xAI team deserves more credit. A visual summary from Artificial Analysis is below:

![](https://substackcdn.com/image/fetch/$s_!NQJl!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd8ac8b4-470f-453a-bf13-56bd05bd27ef_4096x1723.jpeg)

![](https://substackcdn.com/image/fetch/$s_!NQJl!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd8ac8b4-470f-453a-bf13-56bd05bd27ef_4096x1723.jpeg)

这次发布以及其他近期事件,使开放模型与封闭模型之间力量对比的最可能结果发生了重大方向转变。我将逐一展开分析。

This release and other recent events have caused a major change in direction for the most likely outcomes in the balance between open and closed models. I’ll unpack them individually.

从很多方面看,这感觉像是一个新时代的开端。一个竞争更加激烈、但也更需要协调的时代,因为我们正在世界各地部署极其强大的技术。

In many ways, it feels like the start of a new era. An era with much more competition, but also a much higher need for coordination, as we roll out incredibly powerful technologies around the world.

1. 中国重新致力于开源 AI——展现对近期风险的不同解读 1. China’s recommits to open-source AI – showing a different read on near-term risks

许多人直到最近才开始关注中国的 AI 领域,因此他们可能得出结论:公开释放模型是他们的核心策略。实际上,我认为大多数实验室的核心策略与 Anthropic 或 OpenAI 更接近——构建尽可能最好的智能。如今我已经关注并参与中国实验室多年,对他们最初转向将模型开源的最佳解释是实用性。他们需要公开释放模型以获得采用、关注和反馈(尤其是在高价值的湾区市场)。

Many people started following China's AI scene relatively recently, so they may conclude that releasing models openly is their core strategy. In fact, I think most labs have a core strategy far closer to Anthropic or OpenAI: build the best intelligence possible. Having followed and engaged with Chinese labs for years now, the best explanation for their original turn to open-sourcing models is practicality. They needed to release models openly to gain adoption, attention, and feedback, especially in the high-value Bay Area market.

长期以来,中国在解释开源 AI 的作用以及可能的“国家层面战略”方面政策非常有限。据我所知,此前没有高层领导人公开评论过开源 AI。而就在本周,这一情况发生了变化:习近平在世界人工智能大会(WAIC)上发表主旨演讲,非常直接地将中国 AI 生态系统的未来承诺给开源和全球扩散。这一对现状的承诺,恰好与迄今为止最强的开放权重模型发布同一周,成为现代 AI 早期历史中的明确标志。

For a long time, there has been very limited policy in China explaining the role of open-source AI and what a 'country-level strategy' could be. To my knowledge, no senior leaders had publicly commented on open-source AI. This changed this week, too, as Xi Jinping gave a keynote address at the World AI Conference (WAIC) and very directly committed the future of China's AI ecosystem to open-source and global diffusion. This commitment to the status quo, in the same week as the announcement of the strongest open-weight model to date, is a clear marker in the early history of modern AI.

这时正值关于中国 AI 产业未来许多可能路径的讨论期间——他们会保持开放吗?他们能在规模扩张上跟上美国实验室吗?中国是否有一个增长中的收入市场?在这些问题中,焦点一直放在中国的风险承受能力、公司的变现能力,以及任何与公司停止公开释放其*最佳*模型密切相关的理由。

This comes at a time when many potential paths forward have been discussed for the Chinese AI industry: Will they stay open? Can they keep up with American labs in scaling? Is there a growing revenue market in China? With these questions, the focus has been on China's risk tolerance, companies' ability to monetize, and any closely related reason for a company to stop releasing their *best* models openly.

将习近平的承诺与一个非常强大的模型在时间上联系起来,中国已隐含地评论了其对发布开放权重模型的风险容忍度。就目前而言,这是对中国体系中对强势网络安全能力(或生物危害)等话题所感知风险的一种解读。

By tying Xi's commitment in time to a very strong model, China has implicitly commented on its risk tolerance regarding open-weight model releases. For the time being, it is a read on the perceived risks of issues like strong cybersecurity capabilities (or bio-dangers) within the Chinese system.

最简单的解释是,中国政府确实在密切跟踪模型带来的潜在风险——很可能比美国政府的氛围式监管更具技术深度——一旦评估到风险就会采取行动。最简单的解释是,他们认为当前前沿模型不具有实际风险。

The simplest explanation is that China's government is definitely following potential risks from the models closely—likely with more technical depth than the US government's vibe regulation—and would take action if it measured risk. The simple explanation is that they do not find current frontier models to have meaningful risk.

与此同时,中国的经济决策者认为,推广 AI 应用是有益的,这样他们可以在扩大覆盖面之后,再从这个行业获利(正如中国近年来在汽车、太阳能、先进制造业和许多领域所做的那样)。

At the same time, China’s economic decision makers think that having AI adoption is good, so they can make profits on the industry later—after growing distribution (as China has done for cars, solar, advanced manufacturing, and many areas in recent history).

这些在美国 AI 媒体圈看来可能有些令人震惊,因为那里的媒体已经对 Claude Mythos 模型进行了数月的炒作和恐慌渲染。这种反差恰恰是一个极好的现实提醒——世界并不一致认同我们在美国最常听到的那些关于 AI 的叙事。

These can seem somewhat shocking, in an American AI media landscape that has gone through months of hype and fearmongering over the Claude Mythos model. This surprise should be excellent grounding—the world does not have a unanimous agreement with the narratives about AI that we hear most in the U.S.

2. 开放模型:前沿实验室的经济阿喀琉斯之踵 2. Open models as the economic Achilles heel of frontier labs

在引导人工智能讨论的许多叙述者中,有明显的动机去压低最佳开放 AI 模型的感知能力。Dean Ball——他个人支持开放模型,但现在在 OpenAI 工作——写过一篇被广泛评论的关于 Kimi 的帖子,其中他对开放模型说了以下内容。理解这一说法很重要,因为它聚焦于开放模型在 AI 建设经济层面的作用。Dean 说:[1]

Many of the narrators guiding the discussion on AI have clear incentives to depress the perceived capabilities of the best open AI models. Dean Ball—who is personally supportive of open models but now works at OpenAI—had a widely commented-on post with some reflections on Kimi, where he said the following about open models. It is important to understand the statement, as it focuses on the role of open models in the economic side of the AI buildout. Dean says: [1]

解释为什么开放权重模型是_一种减速主义_,对于理解即将到来的世界秩序很重要。他说得对。

Explaining why open-weight models are _a form of decelerationism_ is important to understanding the coming world order. He is right.

对于前沿实验室而言,开放模型在经济上是减速主义的,这将减缓 AI 的净投资和资本支出(capex)部署。原因在于,强大的开放权重 AI 模型大幅压缩了封闭实验室的利润率潜力。这带来两个影响:第一,AI 实验室可用于再投资未来模型的利润减少;第二,市场认为这些公司的终值更低,从而削弱其未来的融资轮次。这两者共同作用,将推迟最具变革性的 AI 模型到来的时间线,但我认为这些影响不足以阻止 OpenAI 和 Anthropic 成为全球市值最高的几家公司之一。

Open models are economically decelerationist for the frontier labs, which will slow the net investment and capex rollout for AI. This is due to the fact that strong open-weight AI models massively reduce the margin potential for the closed labs. This has two effects. First, the AI labs have fewer profits to reinvest into future models. Second, the market sees the terminal value of these companies as being lower, so they will kneecap future fundraising rounds. Together these effects will slow timelines to the most transformative AI models, but I do not see them as strong enough to stop OpenAI and Anthropic from being a few of the top-valued companies in the world.

对我来说,这些对社会总体上是净收益。开放权重模型之所以是_AI 经济扩散_的加速器,是因为它们降低了在特定性能水平下获取智能的入门价格。开放模型也鼓励定制化。问题在于,这种扩散本质上远比前沿 AI 实验室的产品慢——那些实验室直接向开发者销售工具。开放模型的潜力在于,几乎每个企业都可以用它们来构建特定领域的智能体。这种经济扩散需要极长的时间!我将此描述为:开放权重模型的起点慢得多,但可能是指数级更大的曲线。问题是,如果封闭模型在原始能力上领先太多,这种定制能力就可能变得无关紧要。

These, to me, are a net good for society, as open-weight models are accelerationist for _AI diffusion across the economy_ by lowering the entry price for intelligence at a certain level of performance. Open models also encourage customization. The thing is that this type of diffusion is by nature far slower than the products of frontier AI labs; those labs sell tools used directly by developers. The potential for open models is for nearly every business to use them to craft domain-specific agents. This economic diffusion takes an extremely long time! I have described this as open-weight models being on a much slower-starting but potentially bigger exponential. The problem is that if closed models get too far ahead in raw capabilities, this ability to customize can become moot.

我认为,AI 扩散的增强与 AI 实验室权力集中的减弱相结合,对 AI 转型非常积极。它给了我们更多时间去解决新能力的难题,并让更多利益相关者影响叙事——任何一家公司都很可能在安全地控制世界上最重要的技术方面遇到问题。

I see the combination of increased diffusion and decreased concentration of power in AI labs as very positive for the AI transition. It gives us more time to figure out the hard problems of new capabilities and lets more stakeholders influence the story—any one company is very likely to have issues with controlling the world’s most important technology safely.

当然,在这个世界上,对我而言重要的是,最好的模型仍然由美国公司制造,这将使美国能够控制技术及其价值观的发展轨迹。我也预计情况会如此,因为美国拥有更大的资本市场,愿意投资于人工智能(以及日益增长的利润份额),但这并非理所当然。

It is, of course, important to me in this world for the best models to still be made by the U.S. companies, which will allow the US to control the trajectory of the technology and its values. I also expect this to be the case, as the U.S. has larger capital markets that are willing to invest in AI (and a growing share of profits), but it is not a given.

Interconnects AI 是一份由读者支持的出版物。请考虑成为订阅者。

Interconnects AI is a reader-supported publication. Consider becoming a subscriber.

3. 中国的效率优势 3. China’s efficiency advantage

Kimi 的发布博客中有一些技术细节,证实了我们所感受到的持续模型进步背后那些改进。举个例子:

Kimi's launch blog contains technical details that confirm the sort of improvements that are supplying the consistent model progress we feel. To pick one:

训练效率的提升确实会不断累积。它们将带来模型持续而惊人的进步。

Training efficiency really adds up. They will result in continued, incredible advances for models.

顺便说一句,通过生态系统追踪这一特定创新的路径是一个有趣的例子。Kimi Delta Attention(KDA)是在 Kimi Linear 论文中提出的,它与 OLMo Hybrid 所用的 Gated DeltaNet 相似(那是我在 Ai2 期间参与的最后一个 OLMo 模型)。Qwen 的最新模型已转向相关架构,最近的 Nemotron 模型也同样采用了混合架构(但仍更接近 Mamba,而不是 Gated DeltaNet)。看到这类由学术界大力推动的新架构思想如此迅速地被转化到前沿模型,真是令人赞叹。Gated Delta Networks 于 2024 年底提出,其基础是 Mamba 的思想。到 2026 年年中,它们已出现在前沿模型中。

As an aside, tracing the path of this particular innovation through the ecosystem is an interesting example. Kimi Delta Attention (KDA) was introduced in the Kimi Linear paper, and is similar to the Gated DeltaNet used for OLMo Hybrid (my last OLMo model while at Ai2). Qwen's latest models have switched to a related architecture, and the recent Nemotron models are also hybrid (though still closer to Mamba than to Gated DeltaNet). It is great to see new architecture ideas like these, which were heavily progressed by academia, get so quickly translated into frontier-scale models. Gated Delta Networks were introduced in late 2024, building on ideas from Mamba. By mid-2026, they are in frontier models.

我选择聚焦这个例子,部分是因为 Kimi 团队为模型之间的创新赋予了一个引人注目的数字,但主要是为了给关于中国资源效率的更广泛讨论留出空间。

I chose to focus on this example partly because the Kimi team attached a striking number to innovations between models, but mainly to open space for a broader discussion of China's resource efficiency.

越来越清晰的是,中国实验室的资本效率要高得多。在一个由缩放定律决定智能与有效资本成正比的世界里——资本可用于购买算力、数据和人才——这可能是你们的 AI 产业所能拥有的最大优势。对于为何会如此,有许多可能的解释,例如中国研究人员的薪酬更低,却在 LLM 研究难题上更高效,但我们可能永远无法得到如此具体的原因。

It is becoming clear that Chinese labs are far more capital-efficient. In a world where scaling laws dictate that intelligence is proportional to effective capital which buys compute, data, and talent, that may be the greatest strength your AI industry could ever have. There are many possible explanations for why this is the case, such as Chinese researchers being paid less while being more effective at LLM research challenges, but we will probably never get such specific reasons.

自从我撰写关于中国的笔记以来,我听到越来越多关于中国新兴数据产业的消息(在数据方面,它远不及 Anthropic 数十亿美元的预算),以及中国实验室(通过规避出口管制)获得了可观的训练算力。中国公司并没有同样的推理需求(直到最近,Moonshot AI 不得不暂停其 K3 模型的新订阅,而 API 仍然可用),因此他们可能有更多算力用于训练。这些领域对实验室能力的真实情况有很大影响,但我们对其的度量非常有限。

Since writing my notes on China, I have been hearing more about an emerging data industry in China (far behind Anthropic’s billion-dollar budgets for data) and that Chinese labs have access to meaningful training compute (by skirting export controls). Chinese companies do not have the same inference demand (until recently, as Moonshot AI had to pause new subscriptions for access to their K3 model — while the API is still live), so much more of their compute could go to training. These areas heavily affect the truth about the labs’ capabilities, but we have very limited measurement into them.

现实情况是,这些中国实验室募集的资本比美国 AI 生态系统的任何部分都少几个数量级。最直接的比较对象是 OpenAI 和 Anthropic,它们的公开模型稍好一些。其他公司,如 Google 和 Meta,拥有商业史上最大的现金流,但在模型构建上落后了。至于美国的新兴实验室,形势更具竞争性——Thinking Machines 最近发布了他们的第一个模型 Inkling,它很强,但与 Kimi K3 不在同一级别。

The facts on the ground are that these Chinese labs have raised orders of magnitude less capital than any segment of the American AI ecosystem. The most direct comparisons are to OpenAI and Anthropic, which have slightly better public models. Others, such as Google and Meta, have the largest cash flows in the history of business, and are behind in building models. As for American neolabs, the picture is even more competitive — Thinking Machines recently released their first model, Inkling, which is strong but not in the same class as Kimi K3.

这些拥有更多资源的美国公司仍有可能迎头赶上,但你需要认真权衡我们已有的模型质量公开度量,而不是诉诸希望——希望往往反映偏见。如果中国实验室确实具有潜在优势,他们可以继续利用它来构建比所有竞争对手都更强的模型!许多结果都是可能的,K3 应该会提高大多数人对中国在近期内凭借更高效的训练努力彻底领先 AI 能力的概率——即使这不是你预测的最可能结果。

These American companies, with more resources, could still catch up, but you need to strongly weigh the public measurements we have of model quality and not resort to hope — which often reflects a bias. If Chinese labs do have a latent advantage, they could continue to utilize it to build even stronger models than all the competitors! Many outcomes are plausible, and K3 should increase most people’s probability that China can outright lead in AI capabilities in the near future on the back of more efficient training efforts — even if it’s not your most likely predicted outcome.

资本效率的一个主要贡献因素可能在于方法:美国实验室正在投入大量精力,以戏剧性的大步推动前沿,而中国实验室更专注于追赶——这种追赶成本更低。正如在蒸馏中,学生模型通常可以超越教师模型(不仅限于中国实验室的对抗式蒸馏),“努力追赶”而不是“发明下一个范式”的方法可能会产生更强的模型。

A big contributor to capital efficiency is likely in the approach: American labs are spending significant energy pushing the frontier in dramatic, large steps, while Chinese labs are more focused on catching up — and this catch-up is cheaper. Just as in distillation the student model can generally outperform the teacher (not limited to the adversarial distillation of the Chinese labs), an approach of “trying to catch up” rather than “invent the next paradigm” could lead to stronger models.

4. 不断壮大的前沿开放模型生态 4. A growing ecosystem of frontier, open models

在 Kimi K3 发布后的那个周末,我在撰写本文并广泛讨论相关事件时,阿里巴巴宣布即将推出一个 2.4 万亿参数的 Qwen 3.8 模型,并开放权重。历史上,阿里巴巴一直通过其云业务将其最大的模型作为仅提供 API 的产品,因此这又是一次重大的氛围转变,为开放模型经济的下一个篇章打开了大门。即使该模型在基准测试上落后于 Kimi K3,这也表明中国公司可能不仅是在维持其开放模型战略的现状,而是在进一步加大投入。

The weekend after the Kimi K3 release, while writing this and discussing the events broadly, Alibaba announced that a 2.4 trillion parameter Qwen 3.8 model is coming soon with open-weights. Historically, Alibaba has kept their largest models as API-only offerings via their cloud business, so this is another big vibe shift opening the doors to the next chapter of the open model economy. Even if the model is behind Kimi K3 on benchmarks, it signifies that Chinese companies may not only be maintaining the status quo for their open model strategy, but leaning further into it.

如果这个 Qwen 3.8 模型很快发布(即在下一个 Gemini 模型之前),它可能会将谷歌推至拥有最智能模型的实验室排行榜上的第 8 位——而中国在这份榜单上的排名一直在攀升。还有其他关于中国即将推出更多强大模型的传闻,DeepSeek V4 预计将走出其“预览”版本。

If this Qwen 3.8 model releases soon, i.e. before the next Gemini model, it could push Google to the 8th position on the leaderboard of labs with the smartest models – a list that China has been climbing. There are other rumors of more strong Chinese models soon, with DeepSeek V4 expected to graduate out of its “preview” version.

5. 前沿开放权重政策漫长故事的开端 5. The very beginning of a long story of frontier open-weight policy

我认为,如果今天将 Claude Mythos 作为开放权重模型发布,负面结果会相对较小。这是一个有些难以持有的观点,因为我们拥有的公开网络安全评估非常有限,而且这是一个复杂的生态系统(也因为我相信 Anthropic 的许多人)。我仍然坚持这一观点。风险被过度夸大了。

I think that if Claude Mythos was released as an open-weight model today, the negative outcomes would be relatively minor. This is a somewhat challenging opinion to hold, as we have very limited public cybersecurity evaluations and it is a complicated ecosystem (and because I trust many people at Anthropic). I still stand by it. The risks have been over-hyped.

问题在于,对于最强的 AI 模型来说,情况并非总是如此。更强大的模型即将到来——随之而来的是更大的风险——因此,最好的模型以受控、封闭的方式比类似的开放权重模型提前几个月被访问,是一个非常安全的均衡状态。

The problem is that this will not always be the case for the strongest AI models. Far stronger models are coming — and with them increased risks — so it is an incredibly safe equilibrium for the best models to be accessed in a controlled, closed manner several months ahead of similar open-weight models.

用户高度可控的开放权重模型将不断涌现——你无法有效地禁止数字产品,尤其是对恶意行为者而言——因为人工智能训练已被证明在全球范围内可获得的时间比许多分析师预期的更长。

Open-weight models that are very controllable by the user will always be coming — you cannot effectively ban digital products, especially from bad actors — as AI training has proven to be globally accessible for longer than many analysts expected.

尽管如此,在我写这篇文章的时候,美国政府仍在考虑更多旨在限制开放权重模型的措施。最新消息来自 Axios:

Still, as I write this, the government continues to flirt with more measures aimed at restricting open-weight models in the U.S. The latest is from Axios:

这将使美国处于一种非常不对称的状态:美国最好的模型在网络安全任务上有防护栏,但全球参与者可以访问优秀的中国开放权重模型来探测我们的防御。这是禁止开放权重模型不仅损害人工智能自由市场,而且在短期内使生态系统更不安全的众多例子之一。还有其他非常糟糕的后果,例如减缓人工智能应用和人工智能研究的扩散,正如我上面所讨论的。

This would leave the U.S. in a very asymmetric state where the best models in the U.S. have guardrails on cybersecurity tasks, but global actors have access to great Chinese open-weight models to probe our defenses. This is one of many examples where banning open-weight models not only harms the free markets of AI but also makes the ecosystem less safe in the short term. There are other very bad outcomes, such as slowing the diffusion of AI applications and AI research, as I discussed above.

这些平衡很难维持,尤其是当 AI 工具加速模型的进步时,但保持开放与封闭之间的这种现状很重要。一个在能力上真正独处前沿的模型——就像 Mythos 发布时那样——同时还是开放权重的,这将随着我们进入未知的能力领域而带来严重风险。模型将进步得非常快,衡量其全部能力也变得越来越困难。

These equilibriums are very hard to maintain, especially as AI tools accelerate progress in the models, but it is important to maintain this status quo between open and closed. Having a model that is truly alone at the frontier in capabilities — something like Mythos when it was announced — also being open-weight poses serious risks as we go into the unknown of capabilities. Models are going to progress very fast and it is increasingly hard to measure their total capabilities.

我们于是陷入一个在开放模型上走钢丝的世界。不希望如此强大的东西瞬间在全球扩散是合理的,但与此同时,模型的制造者有动力夸大其能力,而他们的竞争对手则有动力夸大其风险。这一切都归结为仔细的衡量和主动加固社会以应对风险向量。

We are then stuck in a world where we are trying to thread the needle on open models. It’s reasonable to not want something so powerful to be diffused globally in an instant, but meanwhile the makers of the models are incentivized to hype their capabilities, and their competitors are incentivized to hype their risks. It all comes down to careful measurement and proactive hardening of society to risk vectors.

这列摇摇晃晃的政策辩论、模型发布和喧嚣反应之车从今天起只会继续前行。自去年 12 月 Claude Opus 4.5 发布以来,我们一直处于快速进步的列车之上,所有关于 AI 如何发展的关键思想都在这条智能体式路径上得到检验。做出正确决策的关键在于评估能力,并独立于那些拥有最大财务利益的公司。当我们进入 AGI 时代的 AI 治理时,所需的众多行动之一是以“曲速行动”(Operation Warp Speed)式的方法来提升国家能力(以及其他独立行动者),使其能够准确评估模型并研究新兴风险。

This careening train of policy debates, model releases, and raucous reactions is only going to continue from today. We’ve been on a train of rapid progress, where all the key ideas of how AI should play out are tested, since the release of Claude Opus 4.5 last December, which sent us down the agentic pathway. The key to making good decisions here is evaluation capabilities, independent of the companies with the largest financial stakes. One of many actions needed then, as we enter the AGI era of AI governance, is an Operation Warp Speed style approach of bootstrapping state capacity (and other independent actors) that can evaluate models accurately, and study emerging risks.

结论:警钟 Conclusion: The wake-up call

开放权重模型通过加速能力的扩散,是 AI 潜在益处和潜在风险的大规模升级。目前,前沿语言模型的负面影响在很大程度上还只是假设,但这种情况不会一直如此。

Open-weight models, by accelerating the diffusion of capabilities, are a massive escalation in the good and the potential bad of AI. For now, the bad side of frontier language models has been largely hypothetical, but that will not always remain the case.

让开放权重模型稍微落后于封闭前沿,是我们缓解风险的自然缓冲。关键在于,我们必须集体行动,在潜在危害出现时加以缓解,而无论开放权重模型落后 3 个月、6 个月还是 9 个月,那都是非常短的时间线。如果我们对开放权重模型进行严厉监管,我担心世界上的许多人会被麻痹,认为我们不再需要采取行动。我们充其量只是稍微推迟了不可避免的事情——开放模型最终将跨越所有关键能力阈值,而且不论法律如何。

Having open-weight models be slightly behind the closed frontier is our natural buffer to mitigate the risks. The key point is that we must collectively act to mitigate potential harms as they appear, and whether open-weight models are 3 or 6 or 9 months behind, that is still a very short timeline. If we regulate open-weight models heavy-handedly, I suspect much of the world will be lulled into thinking we no longer need to act. All we would’ve done is slightly delayed the inevitable — open models will continue to cross all the key capability thresholds eventually and regardless of legality.

理解并受益于这种开放与封闭的共舞,必须成为 AI 社区在今后几年中跨越所有权力和影响力领域的集体行动。

Understanding and benefiting from this open-closed dance must be a collective action from the AI community across all sectors of power and influence over the coming years.

Kimi K3 是一个分水岭时刻,因为前沿开放权重模型如今已成为现实。许多关于开放权重模型风险真正落点的假设将受到检验——我怀疑风险范围会比许多人预期的更窄,而且许多 AI 风险仍将通过封闭且“更安全”的 API 扩散。这些风险的评估将随着 AI 在我们经济中的整合加速而不断发展。我们不能只取其一而不取其二,我们将继续两者兼得。

Kimi K3 is a watershed moment because frontier open-weight models are now real. Many hypotheses will be tested on where risks of open-weight models truly land — I suspect it’ll be narrower than many expect, and many risks of AI will still be proliferated by closed and “safer” APIs. The evaluation of these risks will evolve in time with an acceleration of AI’s integration in our economy. We cannot get one without the other, and we will continue to get both.

因此,2025 年开放模型开始受到更多重视——尤其是在中国以如此明显的领先优势跃居前列之后——人们意识到,世界将不再是单极的,不再只有美国的封闭 AI 实验室决定发展轨迹。2026 年则是此前讨论过的、由真正前沿的开放权重模型带来的潜在风险和加速真正落地之时。

With this, 2025 was when open models started to be taken more seriously — especially when China leaped ahead with such a clear lead — as people realized that it would not be a unipolar world, with only American, closed AI labs determining the trajectory. 2026 is when those previously discussed, potential risks and accelerations due to truly frontier, open-weight models landed.

回应中很大一部分是针对第四点,该点将开放模型的必然结果比作 AI 共产主义,我认为这没有切中要害。具体来说,不加解释地使用“共产主义”一词引发了大量反弹。

A large portion of the response was for the fourth bullet, which compared the inevitable outcome of open models to AI communism, which I think missed the mark. Specifically, the use of the word communism without explanation caused much of the blowback.

互动版:图/公式 + 针对本篇提问 →