开放模型回顾:关于 Kimi K3、Qwen 3.8、习近平 WAIC 演讲、蒸馏、开放与封闭差距以及未来展望的更多内容

Open models recap: more on Kimi K3, Qwen 3.8, Xi's WAIC speech, distillation, the open-closed gap, and what's next

内森·兰伯特 Nathan Lambert · Interconnects · 2026-07-22 · Interconnects ↗

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

本文讨论了开放权重 AI 模型的快速进展,重点关注了 Kimi K3 和 Qwen 3.8 等最新发布,及其对 AI 格局的影响。作者 Nathan Lambert 和 Florian Brand 分析了开放与封闭模型之间的性能差距,认为基准测试常常因评估方法不同而误导对这一差距的认知。他们强调了中国模型的惊人质量,并将其归因于资本效率以及在计算、数据和人才方面的战略投资。讨论涵盖了大型模型后训练的挑战、蒸馏在模型开发中的作用,以及推动开源采用的地缘政治和经济因素。作者总结道,尽管开放模型在编程等特定任务上正在缩小差距,但在长尾能力上仍显不足,且生态系统正迅速专业化以支持这些更大的模型。他们预测开放模型的发布将继续加速,并强调细致评估优于简单比较的重要性。

This article discusses the rapid advancements in open-weight AI models, focusing on recent releases like Kimi K3 and Qwen 3.8, and their implications for the AI landscape. The authors, Nathan Lambert and Florian Brand, analyze the performance gap between open and closed models, arguing that benchmarks often misrepresent this gap due to varying evaluation methods. They highlight the surprising quality of Chinese models, attributing it to capital efficiency and strategic investments in compute, data, and talent. The discussion covers the challenges of post-training large models, the role of distillation in model development, and the geopolitical and economic factors driving open-source adoption. The authors conclude that while open models are closing the gap in specific tasks like coding, they still lag in long-tail capabilities, and the ecosystem is rapidly professionalizing to support these larger models. They predict continued acceleration in open model releases and emphasize the importance of nuanced evaluation over simplistic comparisons.

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 3)

全文 · Full text(逐段中英对照)

开放模型回顾:更多关于 Kimi K3、Qwen 3.8、习近平 WAIC 演讲、蒸馏、开放与封闭差距以及未来展望 Open models recap: more on Kimi K3, Qwen 3.8, Xi's WAIC speech, distillation, the open-closed gap, and what's next

激动人心的消息!我的书旨在与世界分享后训练知识,现已完成并即将发货。您可以在 Manning 或 Amazon 上订购。感谢您的支持。它目前在亚马逊上排名第一的 AI 书籍 :)

Exciting news! My book, which aims to share post-training knowledge with the world, is finished and will be shipping soon. You can order it on Manning or Amazon. Thank you for your support. It is currently the #1 AI book on Amazon :)

Nathan 和 Florian 坐下来讨论开放模型的所有动态。继上周 Kimi K3 发布后,感觉一切都在加速——美中地缘政治、开放与封闭模型的经济学、AI 前沿的安全问题等等。

Nathan and Florian sit down to discuss everything happening with open models. Following the Kimi K3 release last week, it feels like everything is accelerating — the geopolitics of US vs. China, the economics of open vs. closed models, security at the frontier of AI, and so on.

12:47 中国模型为何如此出色?

12:47 How are the Chinese models this good?

17:41 数据、环境以及中国实验室之旅

17:41 Data, environments, and a tour of the Chinese labs

19:47 中国提供商概览:Qwen、DeepSeek、MiniMax……

19:47 Roundup of Chinese providers: Qwen, DeepSeek, MiniMax…

30:25 前沿与近前沿,以及反对禁令的网络安全理由

30:25 Frontier vs. near-frontier, and the cybersecurity case against bans

34:58 蒸馏与 Ben Thompson 之争

34:58 Distillation and the Ben Thompson debate

44:12 预测与前沿模型分级列表

44:12 Predictions and a frontier tier list

可在 Apple Podcasts、Spotify 及任何播客平台收听。其他 Interconnects 访谈请点击此处。

Listen on Apple Podcasts, Spotify, and wherever you get your podcasts. For other Interconnects interviews, go here.

更多教育性后训练视频,请参见我正在制作的课程。

For more educational post-training videos, see the course I’m putting together.

文字记录 Transcript

00:00:06 Nathan Lambert: 好的,欢迎回到 Interconnects。我们正在进行季度开源模型综述,这主要是我们在调侃或解释——不是调侃——为什么这么多蒸馏观点是糟糕的,并理解当前状况。我想上周四是 Kimi K3 发布的时间。我认为在不久的将来我们会看到更多。这似乎不可避免。比如周末,习主席发表了讲话,他直接承诺将开放和开源作为战略。这并不是一个详细的布局状况。

00:00:06 Nathan Lambert: Okay, welcome back to Interconnects. We're doing our quarterly open model roundup, which is mostly us just making fun of or explaining, not making fun, why so many distillation takes are bad and understanding the state of where things stand. I think last Thursday was when Kimi K3 was released. I think we will see much, much more in the near future. It seems pretty inevitable. Like over the weekend, Xi gave his speech where he directly committed to openness and open source as a strategy. It wasn't a detailed layout state of affairs.

Qwen 宣布他们的下一个大模型将是开放权重的,这是一个巨大的变化。我觉得有很多内容要讨论。我想 Flo 你已经开始讨论一些性能差距和蒸馏观点了。所以我们可能可以从那里开始,然后随着我进行,我有一个小列表,我们可以随时浏览我写的博客中的主题,这些都非常细致。所以我认为我们有无限的话题可谈。所以继续你的吐槽吧。

Qwen announced their next big model is going to be open weight, which is a big change of things. I think there's just so much to get into. I think Flo you kind of were already going off on some of the performance gap and distillation takes. So we could probably start there and then as I go I have a little bit a little list and we could always go through the topics and the blog that I wrote which all are very nuanced. So I think we have infinite to talk about. So continue rant kind.

00:01:17 Florian Brand: 是的,我认为,每次模型发布,至少每次开源模型发布,最大的事情是它落后于封闭前沿多少个月。人们喜欢给这个数字一个明确的定义,这真的非常模糊,因为现在我们有这么多不同的基准提供者,以及这么多不同的基准,以至于每个网站——我也不例外——都会拿出他们最喜欢的基准来展示当前模型或新发布模型处于前沿,然后另一方会拿出另一个基准来反驳,显示它实际上落后一年或什么的。很多问题似乎都取决于这个:开源模型落后多少个月。

00:01:17 Florian Brand: Yeah, I think, or the biggest thing at every model release at least at every open model release is how much or how many months it is behind the closed frontier. Um and people love to put a definite uh definitive number onto this uh which is really really mudding because we have so many different benchmark providers these days and such uh so many different benchmarks as well that every site and I'm not innocent in that either um pulls up their favorite benchmarks to show that the current model or the newly released model is at the frontier which is then counted by the other side pulling up another benchmark and showing oh it's actually a year behind or something. Um and it like a lot of it seemingly hinges on that question how many months we open models are behind.

00:02:26 Nathan Lambert: 是的。所以我的挑衅是,一些基准实际上与人们正在做的事情合理相关,这是智能体式编码和智能体式计算机使用任务,而一些基准与长尾相关,我认为 Claude 和 GPT 在那里非常有价值。但如果是这样,比如现在 Claude Code 和 Codex 的市场是什么,如果是软件工程,那么模型在这方面落后几个月可能是一件非常非常大的事情。然后我怀疑这个模型会没问题——免责声明:模型权重还没出来,据说在 7 月 27 日——很多讨论将基于它们会出来的假设。

00:02:26 Nathan Lambert: Yeah. So I my provocation is that some of the benchmarks are actually reasonably correlated with what people are doing and this is agentic coding and agentic computer use tasks and some of the benchmarks are correlated with the long tail which is where I think Claude and GPT is so valuable. But if it's it's like what is the market for Claude Code and Codex right now and if it is software engineering then like the fact that the models are say a couple months behind on that can be a very, very big deal and then I suspect that this model will be okay disclaimer the model weights aren't out yet supposedly on 20 July 27th and a lot of the discussion will impinge on the assumption that they come.

但是人们可以对这种模型进行后训练,很可能在许多人们想要的这类小众领域匹配 Opus 和 GPT。我认为观察——我的意思是我们对开源模型后训练行业有不同的看法,但在使这些模型针对特定高价值任务进行微调方面,有大量的兴奋和进展。这在历史上是在 Qwen 和 GLM 的混合上完成的,GLM 5.2 真正加速了这一点。我很好奇第一个发布博客文章的人,比如“我们在我们的任务上微调了 Kimi K3”,因为我打赌你能获得巨大的收益。我认为即使你比我更多地使用 Kimi K3,但我的直觉是,由于它的规模扩大,后训练可能会有点粗糙,这通常意味着还有很多性能可以从中提取。不。

But like people could post-train this model to very likely match Opus and GPT in many of these kind of niche domains that people want I think watching I mean we both have different views into the post-training open model industry, but there is a ton ton of excitement in progress on making these models like fine-tuned for specific high-value tasks and this has been historically done on a mix of like Qwen and GLM and GLM 5.2 really accelerated this and I I curious on the first person that puts out a blog post like we fine-tuned Kimi K3 on our task because I bet you could get big gains. I think even the you use Kimi K3 more than I do, but my hunch is that it would be a bit of a um rough edged post-training just by how big of a scale up it is and that normally means there's a lot of performance that could still be extracted from it. No.

00:04:03 Florian Brand: 是的。运行、运行、运行,尤其是后训练,那将非常困难,因为你需要一个 B300 节点才能加载权重,这在规模上非常疯狂。所以可能需要一些时间,我听说需要大量的工程工作才能真正让它达到可微调的状态。但是人们你想谈谈使用这个模型,比如你实际上报名了编程项目并使用了它。所以把它发布出来是很好的背景。

00:04:03 Florian Brand: Yeah. Running, running, running and especially post-training that one will be super hard because you need like one node of B300s to just load the weights which is crazy in terms of scale. So will probably take some time and a lot of engineering I’ve heard to actually get this into a state where it’s fine-tunable. But people you want to talk about using the model like you actually signed up for the coding program and used it. So like getting this out there is good context.

00:04:38 Nathan Lambert: 是的。所以我在发布后的第二天就注册了 200 美元的套餐,这是他们最大的套餐,和其他公司类似,但他们还有 40 美元和 100 美元的。但最大的套餐有 100 万上下文,我认为或者至少感觉在 API 请求方面也有一定的优先级,因为很多人在网上说他们经常遇到 API 错误,到目前为止我可以说我过得还不错。在模型能力方面,除了前端它非常好之外,在某些方面它确实表现出色。

00:04:38 Nathan Lambert: Yeah. So I signed up on day after release or so for the $200 plan which is their biggest one similar to all the others but they have like I think $40 and $100 as well. But the biggest plan has 1 million context and I think or at least it feels like it has also some priority in terms of the API requests because so many people online are saying that they hit API errors constantly and so far I’ve been pretty well off if I’m going to say that. And in terms of model capabilities, it at some ways aside from front end where it is really good, it in some ways it really shines and it excels.

即使是我的期望,即使是一些研究任务,比如我在 interconnects 上有的,我们现在有超过一年的开放模型数据,我让前沿模型提出一些我们以前没有做过的有趣分析,因为我们自己做分析并发表,我基本上让它们做点新东西,给我惊喜。很多模型或前沿模型,基本上所有模型都会抓住我们做过的事情,重做数据分析部分,然后做一些奇怪的深奥部分。

Even my expectations even with things like some research tasks like I have or at interconnects we now have over a year of data on open models and I ask the frontier models to come up with some interesting analysis which we haven’t done before because we do our own analysis and have this published and I asked them all right do something new and surprise me, basically. And a lot of the models or the frontier models or basically all models latch onto the things we’ve done redo the data analysis part and then do some weird esoteric parts.

Kimi K3 做了一些更有趣的事情,我明确告诉它去抓取 Reddit,然后它找到了一些我甚至没有考虑过的子版块,然后发现例如 Reddit 讨论比下载数据提前一两个月,或者它们发现有趣的模型比下载量通常起飞早一两个月,比如它们都在关注 Qwen,然后人们下载更多 Qwen 模型,这类分析是开创性的,但这是 Kimi 相比所有其他前沿模型让我惊讶的地方。

Kimi K3 did some more interesting things I’ve told it explicitly to scrape Reddit and then it found some subreddits I haven’t even considered and then found out for example that the Reddit discussions are one or two months more recent or they found the interesting models one or two months before the download numbers usually take off like they are all onto Qwen and then the people download more Qwen models like those kind of analysis is groundbreaking but it is something that Kimi surprised me at compared to all the other frontier models.

一个简单的问题,比如你能用它来做你大部分核心工作吗,比如你有各种任务,你倾向于用 Codex 做大部分,我觉得你是个 Codex 用户而不是 Claude 用户,你认为这个模型在维恩图中重叠的百分比是多少,它就能胜任?

A simple question like can you use this for most of the core work you do in terms of like the exp you have a distribution of stuff you tend to do most of them are with Codex I think you’re a Codex person rather than Claude person like what percentage do you think the Venn diagram overlaps where this model would be fine

00:07:24 Florian Brand: 嗯,这真的取决于我给它多少自由度。比如,我现在在 Prime Intellect 做框架,看到 Kimi K3 的一个大问题是,它的代码简单很多,可读性更好,但缺少一些 Codex 能做到的东西——比如我们说的 56、55,尤其是 54 的水平。所以,我认为 Kimi K3 在这些任务上大概是 54-55 的水平。

00:07:24 Florian Brand: Uh, it really depends on how much leeway I give it. Like, the big thing I have seen with Kimi K3 right now—I'm working on the framework we are doing at Prime Intellect, where I work—and the main thing I found with Kimi is its code is a lot simpler, which makes it way more readable, but it misses some things that Codex just... or like we're talking 56, 55, and especially 54 would be on those levels. So, I would say that Kimi K3 is like 54-55 level for these kinds of tasks.

但如果我读代码,觉得“这代码真不错”,然后用 Codex 过一遍,它会发现所有那些它不擅长的边缘情况,但对于监督运行或跑一些实验来说,它其实非常可用。而且对于一些其他边缘任务,你可以直接让它跑。唯一的缺点是——但这也是因为 API 用户太多,而且服务器在中国——墙钟时间比 GPT 高很多。但我想说,如果我要在日常工作中使用它,我会慢一些,但不会慢到让我觉得“这没法用”。

But if I read the code and I say, "All right, that's really good code," and then I give it a pass over with Codex, and it finds all these niche cases where it doesn't excel, but for supervising runs or for running some experiments, it is actually really usable. And for some other niche things, you can just let it run. The one downside is—but that's also because the API is completely swamped in terms of users, and their servers are in China—the wall clock time is significantly higher than GPT. But I would say if I was to push it and use it in my daily workflow, I would be slower, but I wouldn't be slowed down by so much that I would say, "All right, that's unusable."

00:08:53 Nathan Lambert: 那这和 GLM 5.2 相比如何?因为在我看来,GLM 5.2 的故事还在展开,比如我去旧金山转转,人们会说“我确实用它来做智能体式编码或工作流的一部分”。你怎么看?我觉得你当时是不是也在用 GLM?

00:08:53 Nathan Lambert: And how does this compare to GLM 5.2? Because GLM 5.2 was still a story unfolding in my opinion, where like I would go bop around SF and people are like, "Yeah, I genuinely use this for this part of my agentic coding and/or workflow." How do you—I feel like were you in that camp using GLM at all?

00:09:20 Florian Brand: 是的,我也用过并且现在还在用 GLM,主要是因为我们有一个内部端点,速度很快,或者在那之前,我也用过 API,速度大概每秒 200 或 300 个 token。如果你能以足够好的水平快速完成很多任务,你就会直接用那个模型,而不是去 Codex,然后选择较小的模型,再选择正确的推理努力,再选择快速模式——我就直接用 GLM,得到相同的结果,而且效果很好。它的能力绝对接近 Sonnet 的水平,而且对于很多清理任务,或者只是苦力活,它真的很管用。我觉得你可以用 Kimi K3 作为主智能体,GLM 作为子智能体,来完成很多工作。

00:09:20 Florian Brand: Yeah, I also used and use GLM mostly because we have an internal endpoint which is really fast, and we have—or before that, I also used an API which had, I don't know, 200 or 300 tokens per second. And if you can do a lot of tasks at a good enough level really fast, you just use that model compared to going to Codex, then selecting the lesser model, then selecting the right reasoning effort, then selecting fast—like I just use GLM, get the same result, and it's pretty fine. It definitely is Sonnet-ish level in terms of capabilities, and for a lot of cleanup tasks, for tasks that just are grunt work, it really works. I would say you could probably go really far for a lot of the work with Kimi K3 as the main agent and GLM for sub-agent work.

00:10:24 Nathan Lambert: Kimi 的发布和模型规模有些不同。我认为这些开源模型需要更长时间才能真正优化并在推理提供商中可用。比如 GLM 5.2 很快,但第一,我们还没有权重;第二,我认为它的采用速度不会像 500B、700B 的 MoE 那样快。那里会有更多问题,这是一个非常不同的情况,过去中国模型完成强化学习运行后,会在几小时到几天或一周内发布开放权重,然后生态系统立刻就知道怎么做了。

00:10:24 Nathan Lambert: Something that's pretty different with Kimi's announcement and the scale of models this is. I think it'll take a bit longer for these open models to really be optimized and available across the inference providers. Like GLM 5.2 is pretty fast, but one, we don't have the weights yet, and then two, I don't think it's going to be as fast of a roll out on adoption as the like 500B, 700B MoE. There's going to be more problems there, which is a very different regime where in the past the Chinese models would finish their RL run and release the model with open weights within hours to days maybe a week, and then immediately the ecosystem kind of knew how to do this.

我认为在下一代开放权重模型上,基础设施方面的提升要大得多,我们必须考虑到这一点,就像封闭实验室在发布模型之前会在幕后做这些工作。所以这有点像在操纵时间差,可能会导致人们真正能够对 Kimi 进行后训练并大规模用于工作流程的时间再推迟一个月。作为开放权重的爱好者,我们喜欢说,只有当封闭模型可用时,你才能利用时间差,但现在开放模型中也出现了类似的动态,比如 Kimi 的 API 完全崩溃了。供应太多,需求太多,供应不足。所以这个模型并没有立即扩散,我只是在思考这与性能时间差的关系。

I think there’s a lot bigger of an infrastructure kind of uplift on this next scale of open weight models which I think we have to factor in like the closed labs do this behind the scenes before announcing the models. So it’s just like that is kind of manipulating the time gap in a way where it could be like an extra month before people can actually post-train and use Kimi at scale for their workflows. And like we love to say as an open weight fan like oh it’s only when the closed model is available that you could take the time gap but like now there’s similar dynamics in open models where it’s like the Kimi API is totally broken. There’s too much supply. There’s too much demand. There’s not enough supply. So it’s not like this model is immediately diffusing like the I’m just I’m just thinking about this as it relates to the performance time gap

00:11:54 Florian Brand: 因为这是事实,但另一方面,开放生态系统在过去几个月里已经相当专业化了。在你们最初的发布中,他们都会有一些合作伙伴提前获得权重。他们提前几天甚至几周就发布了 vLLM 补丁,这与一年前完全不同,那时权重一发布,模型制造商就说“好吧,你们自己搞定”。所以我预计第一天的普遍可用性会相当不错,然后所有提供商开始竞争优化,以获得越来越高的速度,因为这是很大的声望。

00:11:54 Florian Brand: because that that is true but on the other hand the open ecosystem has professionalized quite a lot in the last few months. uh like during or in your initial roll out they all come with some partners which have the weights beforehand. They have the vLLM patches out days or or even weeks before these days which is completely different from from a year ago where basically weights got dropped and the model makers were like all right you got to figure this out. So I expect like the general availability on day one will be pretty okay and then race starts of all the providers starting to optimize to get even higher and higher speeds because it’s so much prestige.

00:12:47 Nathan Lambert: 是的。好的,有两个方向要讨论。为什么我们认为中国模型能这么好?我想我写过关于我们 Discord 中与 Epoch 的 JSD 的辩论,我认为那很好,我在我的文章中有这一部分,我逐渐认为中国实验室在资本效率上更高,你可以将资本转化为算力、数据和人才,从而使模型更好,我认为这非常重要,如果这真的是某种结构性优势的话。无论原因是什么,我认为原因可能是人才受过更好的训练,他们的教育体系是为了解决那些能让 LLM 更好的问题。

00:12:47 Nathan Lambert: Yeah. Okay. Two directions to go. Why do we think the Chinese models are able to be this good? I think I’ve wrote about there’s a debate in our Discord with in with JSD at at Epoch and I think it’s very good and I had this section in my piece that I’m like coming around to think that the Chinese labs are more capital efficient and you can turn capital into compute data and talent in a way that makes the models better and this is really I think this is super important if it actually is some structural advantage whatever the cause I think the cause could be talent is better trained for whatever their education system was to work on problems that make LLMs better.

也可能只是中国的算力、人才和所有东西的成本都更低,无论是补贴还是平均工资较低。但这是一个非常重要的问题,随着我们不断迭代模型。如果下一代模型对 Anthropic 来说要花费 100 亿美元,而对 Kimi 来说只要 40 亿美元,这可能非常巨大,但原因尚不清楚。例如,我认为 Kimi 的工程师 Big Eagle 回复了我的推文,他说这有帮助,因为我们不是试图推动前沿,我们只是试图追赶,这真的可能是一种心态问题,即中国实验室的目标范围设定方式使得他们构建这些模型的成本低得多。

It could just be that all the compute and talent and everything cost way less in China somehow. Whether it’s a subsidy, whether it’s just average pay being lower. But this is a very big deal as we turn the crank in the model iterations. And if a next generation model costs $10 billion for Anthropic but only $4 billion for Kimi like this this is like could be very huge but it’s not clear why this is the case. For example, I think Big Eagle the Kimi engineer replied to my tweet on this and was like it helps because we’re not trying to push the frontier. are just trying to catch up, which really could be a mindset thing where how the the goals of the labs are scoped in in China so that it cost them way less money to build these models.

但在过去一年里,我们问了很多问题,比如中国模型是否会落后。我曾认为,由于训练的资本密集性,封闭和开放模型之间的差距会扩大,但似乎情况正相反,这很难解释,但你同意实验室比我们预期的更能跟上吗,特别是中国实验室,为什么?

But I in the last year we’ve asked a lot of questions on like will the Chinese models fall off. I have thought that the gap between closed and open bottles would grow due to this kind of capital intensity of training and it seems like it’s going the opposite direction which is just like it’s it’s hard to unpack but like do you agree that the labs are keeping up a bit more than we would have expected as in the Chinese labs and why?

00:14:43 Florian Brand:嗯,我实际上根据我们去年的回顾,看了我们今年的预测,我们基本上说差距会保持在几个月之内。所以这个预测似乎大体上成立。幸运的是,我们没有给出具体数字,是 3 个月、6 个月还是 9 个月,所以我们在那方面是安全的。但我认为我们在中国和这些人交谈时,我们俩都感觉到的一般情况是,研究人员本身是两三百人的团队,都是二十多岁,都只想把一个模型做得非常好。他们似乎不做任何支线任务。

00:14:43 Florian Brand: Well, I actually looked at our predictions for this year based on our last year's recap, and we basically said that the gap will stay within a few months. So that prediction seems to largely hold. Luckily for us, we didn't put a concrete number on whether it's 3 months, 6 months, or 9 months, so we are safe on that side. But I think the general thing we both felt when we were in China and talking to these people is that the researchers themselves are teams of two or three hundred people, all in their mid-20s, and all just want one model to be really good. They don't seem to do any side quests.

他们似乎不做任何偏离这些事情的事情。而且在算力方面,这对我们来说是一个很难回答的问题,尤其是随着这些中国芯片现在开始投入使用。我们还有……我也认为芯片走私在过去六到九个月里大幅增加,或者说被走私的芯片已经开始上线。

They don't seem to do anything that deviates from these things. And in terms of compute, which is a really hard question for us to answer, especially as these Chinese chips are now coming online, we have... I also think chip smuggling has increased substantially in the last six to nine months, or the chips that have been smuggled have started to become online.

00:15:58 Nathan Lambert:走私是绕过出口限制的通用术语。如果芯片在马来西亚,而他们在使用,我认为那也算类似情况,而且我认为在过去六到九个月里,这种情况大幅增加。这在一定程度上是那个的结果,而你说……但我只是想提出来:我确实认为他们现在拥有的算力比他们训练上一代模型时要多得多。

00:15:58 Nathan Lambert: Smuggling is a general term for getting around export restrictions. If the chips are in Malaysia and they're using them, I count that similar, and I think that that has massively increased in the last six to nine months. This is partially the result of that, and you're saying... but I just wanted to put that out there: I do think that they have a lot more compute than they did when they were training the previous generation of models.

00:16:24 Florian Brand:是的,就像我们……为了提供背景,两周前我认为 LongCat 发布了他们的模型,他们声称,而且我们知道这很可能是真的,完全在中国芯片上训练。他们没有公开具体是哪些,但人们猜测是华为的一些 Ascend 芯片。随着国内生产加速,它们可能被大量使用,或者用于训练,但它们对推理特别有用,而推理也是训练的重要组成部分。

00:16:24 Florian Brand: Yeah, like we... just for context, two weeks ago I think LongCat released their model, which they claim, and we know that it is very likely true, is trained entirely on Chinese chips. They didn't specify publicly which ones, but people speculate that it's some Ascends from Huawei. As domestic production ramps up, they're probably used most, or they are used for training, but they are especially useful for inference, which is a huge part of training as well.

所以他们可能在训练部分使用 Nvidia 和其他芯片的混合,然后在推理部分使用越来越大的比例,这非常重要。所以我认为他们的整体算力在增加,而且他们实际上没有很多用户。所以他们不需要像 ChatGPT 那样为 10 亿用户提供动力,也不需要像 Anthropic 那样为成百上千的企业提供服务,因为他们没有那么多付费客户。

So they probably use some mix of Nvidia and other chips for the training part, and then an increasingly larger part for the inference part, which is really important. So I think their overall compute is increasing, and also they don't actually have a lot of users. So they don't need to power 1 billion users like ChatGPT has to do, or hundreds or thousands of enterprises like Anthropic has to do, because they don't have that magnitude of paying customers.

00:17:41 Nathan Lambert: 是的。而且我认为即使是那些付费客户,至少在企业端,当你支持这些东西时,也会有公司时间和闲聊。即使你不是研究人员,即使这不是你的工作,它确实会改变公司的注意力。如果 SSI 推出一个好模型,那将是对干扰问题存在的最终验证,但这是题外话,我们可以稍后再谈。我认为数据和环境行业开始出现的迹象也已经有了。

00:17:41 Nathan Lambert: Yeah. And I think even those paying customers also, at least on the enterprise side, there’s just like there is company time and chatter when you’re supporting these things. Even if you’re not a researcher, even if it’s not in your job, it does change the attention of the company. And if SSI comes out with a good model, it’ll be the ultimate validation that distractions are a problem, but that’s an aside that we can wait on. I think there’s also rumblings of the data and environments industry starting to appear there.

你还记得具体有哪些吗?因为我们在中国的时候,有点震惊他们似乎很少利用外部数据。所以就在我们旅行几个月后——我们是四月份去的,然后仅仅几个月后的七月份,我们就听到一些关于中国新公司想要购买数据之类的消息。这是一个有趣的时间线,展示了这种变化。

Do you remember any specific ones? Because when we were in China, it was kind of shocking how little they seemed to utilize external data. So just a few months after our trip—we went in April and then just months later in July—we’re hearing a few things of new companies in China and them wanting to buy data and things. And that is a funny timeline of how that changes.

00:18:40 Florian Brand: 我会对他们实际告诉我们的内容加上误差线。

00:18:40 Florian Brand: And I would put error bars on what they actually told us.

00:18:44 Nathan Lambert: 那是因为时间太近了,我不知道。

00:18:44 Nathan Lambert: And that’s because it’s so close in time that I don’t know.

00:18:49 Florian Brand: 是的,那可能是真的。但这类事情很难准确指出。我想说的是,购买外部数据似乎正成为一个更重要的因素。这将有助于开放模型赶上封闭模型,如果他们只是购买相同的数据,也许还能打折,因为他们购买数据环境的时间较晚。但这确实是一个因素。至于这个因素有多大,我们不知道。我们没有任何公开的见解,而且我怀疑我们不会从任何人那里得到这些见解。所以这绝对是为什么我们能够赶上或提高他们模型分数的一部分原因。

00:18:49 Florian Brand: Yeah. That might be true. But like those things are hard to pinpoint. I would say it seems like the buying of external data is becoming more of a factor. Which will help the open models catch up to the closed ones if they just buy the same data maybe at a discount because they buy the data environments later. But it is a factor. How big of a factor we don’t know. We don’t have any public insights and I doubt that we will get those insights from anyone really. So that’s definitely one of the parts why we are able to catch up or improve their model scores.

00:19:47 Nathan Lambert: 好的,接下来盘点其他中国模型提供商。我们已经讨论过 Kimi、智谱/GLM。我认为很快会有更多非常优秀的 GLM 模型,可能会叫 GLM 5.5 之类的。关于 Qwen,我们讨论过他们即将推出的最大模型。我要说的是,Qwen 的最大模型相对于其小模型的卓越性,在性能的绝对排名上往往没有那么突出,这可能是专注的代价。我认为这与云公司有关。几乎可以说,如果你眯着眼看,它几乎就像谷歌。

00:19:47 Nathan Lambert: Okay, roundup of other Chinese model providers. We've talked about Kimi, we talked about Zhipu / GLM. I think there will be more GLM models soon that are very good. They might call it like GLM 5.5. Um, Qwen, we talked about their biggest model coming. Qwen's biggest models, I will say, have tended to, relative to the excellence of their small models, not have the same absolute ranking in performance, which is probably a cost of focus. I think it goes with cloud companies. It's almost like, if you squint, it's almost like Google.

就像 Qwen,阿里巴巴在这里有如此多的机会,而通过这些小模型让开发者与阿里巴巴 Qwen 关联起来,对他们的云业务来说是一个巨大的机会,我认为他们在这方面非常成功。但他们的最大模型一直不如小模型出色。所以我不期望他们的模型能像 Kimi K3 或 GLM 5.2 那样具有突破性。我预计它会被新闻广泛报道,作为中国开源领域的重要发布,但我认为它不会像新闻故事那样持续。嗯,DeepSeek,你可以随意插话。

It's like Qwen has, Alibaba has so much opportunity here, and the opportunity of getting developers associated with Alibaba Qwen with these small models is such a huge opportunity for their cloud that I think they're succeeding wildly. But their big models have always not been as excellent as their small models. So I don't expect their model to be as breakthrough as Kimi K3 or GLM 5.2. I expect it to be covered in the news as major open source, as the open source name in China drops giant model, but I don't think it will be as sustained as a news story. Um, DeepSeek, you can chime in whatever.

00:21:01 Florian Brand: 有趣的是,不知道你关注了多少,但他们有一个端点可以用于预览版本,而且他们每天更新这个端点,所以他们的迭代周期非常快,因为我们在所有这些 Twitter 基准测试中取得了进展。所以很多 SVG 和 three.js 的东西,比如所有这些视觉生成任务,模型在过去几天里改进很大。所以他们找到了某种快速反馈机制,其他公司也有。我们知道这一点,或者 Cursor 有很多关于他们如何快速迭代的博客。但他们似乎不断上传新的检查点并提供使用。

00:21:01 Florian Brand: The interesting thing is, don't know how much you follow this, but they have an endpoint which you can use for a preview version, and they've updated this endpoint daily, so they have some really fast iteration cycle because we progress in all these Twitter benchmarks. So a lot of these SVG things and three.js, like all these visual generation tasks, the model has been improving a lot over the last few days. So they have figured out some kind of fast feedback mechanism, which other companies have as well. We know this, or Cursor has a lot of blogs about this how they iterate really fast. But they seem to continuously upload new checkpoints and make them available.

00:21:47 Nathan Lambert: 嗯,但我同意。我猜这就像在他们最终的强化学习运行中有一个时间门控。在强化学习运行结束时仍在轻微改进,他们只是在检查框。

00:21:47 Nathan Lambert: Um, but I agree. I'm guessing it's like a time-gated within their final RL run. It's like still slightly improving at the end of their RL run, and they're just like checking the box.

00:22:03 Nathan Lambert: 好的。Qwen DeepSeek V4 应该会推出预览版。嗯,关于 DeepSeek V4,我认为 flash 模型实际上更受欢迎,那是他们较小的模型,似乎是人们的绝对主力。所以我认为那是他们值得关注的模型。我不期望 V4 Pro 会有戏剧性的突破。这类似于如果小米很快发布新的 MiMo Pro 模型。我不期望它有那么大的下降,但它可能会是一个非常扎实的模型。只是很难说。他们仍然是一个相当新的进入者。MiniMax,我认为,在玩不同的游戏。我不认为 MiniMax 在追求 Kimi/GLM 那种登月式的 AGI 氛围。

00:22:03 Nathan Lambert: Okay. Qwen DeepSeek V4 is supposed to come out a preview version. Um, the thing about DeepSeek V4, I think, is that the flash model is actually way more popular, which is their smaller one, which seems to be an absolute workhorse for people. So that I think is the model to watch for them. I don't expect V4 Pro to be a dramatic breakthrough. This is similar to anything like if Xiaomi were to release a new MiMo Pro model soon. I don't expect it to be as big of a drop, but it would probably be a very solid model. It's just like it's hard to know. They're still a pretty new entrance. MiniMax, I think, is playing a different game. I don't think MiniMax is chasing this Kimi/GLM moonshot to AGI type vibe.

00:22:46 Florian Brand: 哦,我会,我会不同意这一点。

00:22:46 Florian Brand: Oh, I would, I would disagree there.

00:22:49 Nathan Lambert: 你认为,你认为 MiniMax 还在其中吗?

00:22:49 Nathan Lambert: You think, Do you think MiniMax is still in this?

00:22:52 Florian Brand: 是的,我我我我认为他们他们正看到这种紧张局势,尤其是因为他们是一家像 GLM 这样的上市公司,如果你看看股票表现,最近几天那些股票真是惨不忍睹,嗯,这似乎产生了巨大的影响,有趣的部分将是嗯许可证,因为他们已经多次更改许可证,嗯,变得越来越严格,嗯,如果在习的讲话之后再次改变主意,嗯,看看 MiniMax 是否会回到完全开放的许可证将会很有趣。另外一件有趣的事情是嗯,K3 将采用哪种许可证,因为他们说过会开源,但我不认为他们对实际采用的许可证做出任何承诺。

00:22:52 Florian Brand: Yeah, I I I I think they they are seeing the tension especially because they are a public company similar to GLM and if you look at the stock performance RIP those stocks in the last few days um it it it it make it seems to make a huge difference and the interesting part will be uh the license because they’ve changed the license a lot uh to be more and more restrictive and um if there’s now a change of heart again after the Xi, uh, speech. Uh it will be interesting to see whether MiniMax goes back to completely open licenses. It’s also an interesting thing to see um which license will be the license for for K3 because they have said they will open source it but I don’t think they have done any commitments in terms of the actual license where you put on top.

00:23:45 Nathan Lambert: 是的。我的意思是那那非常重要。是的,我们我们拭目以待。嗯,Ling、美团、LongCat 有点类似,都是非常强大的模型,可能在内部获得了很多价值。但没有同样的开发者突破。嗯,所以那大概有七八个中国实验室。我可能遗漏了一些。我们也可以谈谈美国实验室。另外嗯,Gemini 3.6 flash 发布了。看起来不错。就像就像就像是一个小小的提升。更快了,话少了,但好像并不重要。我们打算停止,我们将停止分享这个。嗯,这就是 Gemini 在我们这里得到的提及量。

00:23:45 Nathan Lambert: Yeah. I mean that’s it’s super important is the thing. Yeah, we we’ll see. Um, Ling, Meituan, LongCat kind of similar, very strong models, probably getting a lot of value out of them internally. Aren’t don’t have the same developer breakthrough. Um, so what that’s like seven seven to eight Chinese labs. I might have forgotten some. And we can also talk about US labs. Aside um, Gemini 3.6 flash dropped. It looks fine. It’s like it’s like it’s it’s a tiny bump. It’s faster. It’s less of a yapper, but like doesn’t really matter. We’re going to stop we’ll stop sharing this. Um that’s that’s the amount of mention that Gemini gets for us.

00:22:46 Florian Brand: Oh, I would, I would disagree there.

但我确实认为值得稍微谈谈美国生态系统。我认为有一些新兴参与者。Thinking Machines 发布了他们的第一个模型。我和他们中的一些人聊过。他们非常愿意参与如何用 Tinker 制作一个可微调模型的研究,我认为这是大多数开源模型构建者真正推荐的研究领域。我认为如果你能在那里获得心智份额,你将获得大规模采用,因为这更多是关于针对实际任务的可微调性,而不是拥有最好的数字。嗯,所以这是他们的 Inkling 模型,一个一万亿参数的模型,得分不错但不是前沿水平。

But I do think it’s worth talking about the US ecosystem a bit. I think there are emerging players. Thinking Machines released their first model. I’ve talked to some of them. They’re very on board for figuring out how to make a fine-tunable model with Tinker, and I think that’s a research area that I really, really recommend for most of the open model builders. I think if you can get mind share there, you will get massive adoption because it’s more about being fine-tunable for real tasks than it is about having the best numbers. Um, so this was their Inkling model, which is a one-trillion-parameter model that has decent but not frontier scores.

我认为,有点像 DeepSeek V4,他们计划发布一个更小的模型,总参数约为四分之一大小,性能非常好。如果 Inkling Small Preview 在几周内发布,我确实认为那将是一个真正被使用的模型。它的尺寸非常适合自动化任务和特定领域任务,可能不像 Kimi 和 GLM 5.2 那样是通用智能体类型,但我认为这非常适合他们的业务。嗯,我知道还有一些其他——我想说美国较小的参与者似乎——比如 Arcee 今年早些时候发布了他们的模型,仍在继续前进。Poolside 已经开始发布一些模型。

I think, kind of like DeepSeek V4, they’re planning to release a smaller model, which is about a quarter of the size in total parameters, and has really, really good performance. And if Inkling Small Preview comes out in a few weeks, I do think that that will be a really used model. It’s a good size for automating tasks and domain-specific tasks, and might not be a general agent type thing like Kimi and GLM 5.2, but I think that suits their business really well. Um, I know that there are some other—I would say the smaller players in the US seem well—like Arcee released their models earlier this year, still chugging along. Poolside has started releasing some models.

他们在过去几个月里发布了一些模型,并且似乎准备在此基础上发布更多模型。所以他们真的在——Reflection 永远处于“模型即将推出”的状态,他们确实应该发布一些模型或代码或其他东西,以便如果他们真的致力于开源,就能开始启动开发者飞轮。这需要很多——实际上很难把模型发布出来。比如我和 Thinking Machines 的一些人聊过,感觉就像,哦,这实际上需要做很多工作,我想。嗯,Nvidia 也在继续前进。我认为他们现在是稳定的参与者。他们坚持发布模型。他们很快就会发布更多。他们发布了很多数据。我正在催促他们发布 Qwen 风格的小模型,比如 Gemma。

They’ve gotten a few out in the last few months and seem poised to release more models on top of that. So they’re really going—Reflection is perpetually in the model-coming-soon camp, and it really behooves them to get some models or some code or something out so that they can just start getting the developer flywheel going if they’re really committed to open source. It just takes a lot—it’s hard to get the models out. Like I talked to some people at Thinking Machines, and it’s kind of like, oh, that’s a lot of work to actually do this, I think. And um, Nvidia chugging along. I think they’re at the stable player at this point. They’re keeping to release models. They’ll release more soon. They release a lot of data. I’m bullying them to try to get them to release Qwen-style small models, which is like Gemma.

Gemma 只有这些——像 Qwen 的竞争对手模型,非常受欢迎。嗯,Gemma 模型有点——它们在尺寸或架构上各不相同,但 Gemma 模型在采用率上确实与 Qwen 模型非常匹配。嗯,我不确定它们是否易于用于研究,这可能需要一段时间。可能需要多次迭代。就像现在很多语言模型研究都是围绕小型 Qwen 模型和基于 Qwen 的模型设计的,这需要时间。比如人们非常了解如何使用这些模型以及研究结果。所以我希望 Gemma 继续推出,并能在那个细分市场中竞争。我不知道我是否漏掉了谁。

Gemma only has these—like Qwen competitor models that are super popular. Um, the Gemma models are a little—they’re all over the place in sizes or in architectures for the sizes and things like this, but the Gemma models are really, really matching the Qwen models in terms of adoption. Um, I’m not sure they’re as easy to use for research, which could take a while. It could take multiple iterations. Like so much of language model research is now designed around small Qwen models and Qwen-based models that it takes a while. Like people know how to use these models really well and with the research results. So I hope Gemma keeps coming and can kind of compete in that niche. I don’t know anyone that I missed here.

00:27:22 Florian Brand: 不,我认为两者都是大玩家。呃,它正在变得更广泛。呃,就模型创作者而言,比如去年,除了 Gemma 3 和嗯 GPT-OSS 之外,我们还有任何发布吗?

00:27:22 Florian Brand: No, I think both are the big players. Uh, it’s—it is becoming broader. Uh, in terms of model creators, like last year, did we have any release aside from Gemma 3 and um GPT-OSS?

00:27:41 Nathan Lambert: 我认为 GPT-OSS 2 会大放异彩,显然 Nemotron 也是如此。嗯,哦,还有年初的 Llama 4,但我不想让它被遗忘。我们看到越来越多的参与者正在加入,并以惊人的速度推出模型。

00:27:41 Nathan Lambert: was GPT-OSS 2 would go hard and obviously and obviously Nemotron as well. Um, oh, and I think Llama 4 at the start of the year, but uh, I don’t want that to be forgotten, but we are seeing like more players are are are now joining and turning out models at a really incredible rate.

00:27:59 Florian Brand: 比如 Poolside 在过去两三个月里发布了三四个模型。他们似乎找到了某种方法,能够相当稳定地推出模型。这也是我们在开源方面看到的情况。比如 GLM,我认为他们的模型发布迭代时间现在在 1 到 2 个月之间,每次迭代都越来越好,这与闭源实验室的做法非常相似,比如我们现在大约每 6 周就会看到一个新的 GPT 或新的 Claude。因此,就拥有足够好的流程来发布越来越强的模型而言,开源生态系统确实已经找到了方法,或者说看起来已经找到了。

00:27:59 Florian Brand: like Poolside has been releasing three or four models in the last two or three months. Uh and they seem to have figured out some way to turn out models pretty consistently. Um and that’s also something we are seeing on the open source side as well. we are talking about GLM like I think their iterations uh times for the model releases are now between 1 or 2 months with each new iteration becoming better and better which closely resembles what the closed labs are doing like we get a new GPT we get a new Claude every uh 6 weeks or so these days uh so in terms of having uh good enough pipeline uh to release stronger and stronger models they have to or the open source ecosystem has really figured it out or seemingly figured it out.

00:28:59 Nathan Lambert: 是的,我同意。这很有希望,但有趣的是,美国生态系统开始发布一些模型,然后你又有 Xi 和这两个模型。要赶上真的很难,因为训练出人们真正使用的模型需要大量的机构专业知识。我认为这正是美国发布模型的公司现在意识到的:这些不仅仅是刷榜的蒸馏知识产权盗窃模型。

00:28:59 Nathan Lambert: Yeah, I agree. It’s it’s promising, but it is also so funny that like the US ecosystem started releasing some models and then then you have like Xi on the mic and these two models. It’s just like it’s so hard to catch up because it takes a lot of institutional expertise to train models that people actually use. And I think this is is what the American companies that are releasing models are now realizing is like these are not just benchmaxxed distilled IP theft models.

这些是真正优秀的模型,人们会在内部交易基准上进行比较,然后发现要在可衡量的方面击败它们有多难。我认为我从美国一些交易模型的人那里感受到了这种情绪,就是……我认为人们应该在规模和可微调性上创新,并尝试利用这个离自己很近的潜在市场,但每家公司都面临着巨大的压力,要发布一个可以称为“前沿”的模型。我认为投资者对这么多参与者都有这样的期望,他们正在尝试做一件相当困难的事情,未来一年中美平衡如何发展将会很有趣。

These are like genuinely good models that people are comparing to on their internal trading benchmarks and then like seeing how hard it is to beat them on measurable things. And I think that that is like I I’ve I’ve picked this sentiment up from a few people in the US trading models and it is just like there’s some I I think people should innovate on like size and fine-tunability and try to like use this potential market that is really close to home but also the pressures for every company is so high to release a model that you can claim as Frontier. I think investors expect that out of so many of these players that they’re kind of trying to do a a pretty hard thing and it’ll be interesting how the next year unfolds for the US China balance.

00:30:25 Florian Brand: 是的,我认为,总的来说,我和其他很多人都讨论过整个生态系统,这也是你在开头提到的。我认为我们越来越多地看到模型能力之间的分化:对于很多任务来说,比如很多编码任务,当前的前沿模型,无论是开源的还是闭源的,都已经足够好了。改进在这里感觉越来越不重要了。

00:30:25 Florian Brand: Yeah, I think or in general I and a lot of other people have talked about the general ecosystem and that’s also something you’ve talked about at the very beginning. I think we are seeing more and more of a split between the capabilities of models that is good enough for a lot of tasks like uh for a lot of coding tasks the current frontier models both open and closed are good enough. um improvements feel less and less uh important here.

但如果我们看看前沿中的前沿,比如发现新的数学证明、找到新的疗法、开发新药、发明新事物,那似乎是完全不同的挑战,并且很可能在相当长的时间内由最前沿的模型主导。那么关键问题就变成了:这在可寻址市场中有多大影响,以及这会有多受关注。我认为,或者说我的基本判断是,我们正看到前沿越来越封闭。我们在网络安全领域的 Mythos、生物技术领域都看到了这一点——这些模型不会对所有人开放,甚至可能不会对外部合作伙伴开放,如果我们考虑到有报道说 Anthropic 正在建立或创建一些内部实验室来开发药物的话。

But if we look at the frontiers frontier, so finding new math proofs, finding new cures, finding new drugs, and inventing new things, that seems to be a whole different beast and probably will be dominated by the very frontier for quite a long time. The big question then becomes how much does that matter in terms of the addressable market and also how much of a focus will this be. I think, or my general base case is that we are seeing the frontier close down more and more. We have seen this with Mythos for cyber security GPT... or for biotech that those models won’t be accessible for everyone and maybe not even external partners if we consider the reports that Anthropic is now spawning or creating some internal labs to develop drugs.

所以,最前沿的模型对所有人都是不可及的,而接近前沿的能力则变得越来越商品化,这会产生许多不同的影响,尤其是当你考虑到网络安全等问题时。两三天前有来自 Hugging Face 的报道说,他们有一个智能体试图入侵他们的系统。他们尝试用 GPT 和 Claude 来分析,但都失败了,因为所有的防护栏都阻止了它们。所以他们不得不使用 GLM,一个能力较弱的模型,但它对这种防御性行动没有防护栏。他们不得不使用一个较差的模型来防御自己或分析数据,这是一种可怕的状态——美国公司现在依赖较弱的模型,因为封闭的前沿模型对它们不可及。

So the very frontier is inaccessible for everyone and then the near frontier capabilities is becoming more and more commoditized which has a lot of different implications especially if you think about things like cyber security. There was that report from Hugging Face two or three days ago that they had some agent trying to hack their system. And they tried to analyze it with GPT and with Claude but were unable to because all the guardrails blocked them. So they had to use GLM, a lesser capable model, but it had no guardrails for this kind of defensive action. And they had to use a worse model to defend themselves or to analyze the data, which is a horrible state to be in that we have US-based companies now relying on lesser models because the closed frontier is inaccessible to them.

00:33:10 Nathan Lambert:是的。我认为这实际上是不采取任何行动的最佳论据之一。如果世界其他地方都能使用这些开放模型,而我们却禁止美国公司使用它们,那么美国公司的防御能力与世界各地的攻击者之间的差距就会越来越大。我们可以争论在当前能力水平下网络风险有多紧迫,但如果你在结构上设定让防御者不会随时间变得更好,而攻击者可以,那么网络风险何时会变得更加真实?很明显,如果你禁止美国公司使用最好的中国开放权重模型,那就会导致这种局面。

00:33:10 Nathan Lambert: Yeah. And I think this is actually one of the best arguments for not doing anything. It’s like if the rest of the world has access to these open models and we ban them for the companies in the US to use and it’s just like a growing disparity between US companies ability to defend and the attackers all over the world in terms of cyber and we could debate like how much of an immediate risks the cyber stuff is at the current capability levels but if you’re setting it up structurally so that the defenders get don’t get better over time and the attackers can like that seems like when the Why would cyber risk become more real? And that would to be very clear that would be if you ban the best Chinese openweight models from being used at companies in the US.

而这种禁令很可能是一种“影子禁令”,即威胁采取法律行动或惩罚,但又不明确具体的执行路径。现在有很多关于此的讨论。我不确定我们是否有很多要说的,但很明显,华盛顿正在尝试各种方式来限制美国使用最好的中国开放权重模型。我认为这是某些恐惧煽动的下游结果。我们也会过渡到蒸馏问题。所有这些都源于美国主流 AI 媒体叙事,将中国模型描绘成窃取知识产权、危险或与中国政府(一个威权政府)有关联。

And this ban would likely be a kind of shadow ban, which is the threat of legal threat of legal action or punishment without it being clear on exactly what the pathway to do it is. And there are a lot of talks about this right now. I don’t like like I don’t know if we’re going to have a ton to say about this, but it’s clear that DC is flirting with different ways of restricting the best Chinese openweight models in the US. This is I think downstream of some fear-mongering. We’ll transition into the distillation question too. It’s like all these things from the primary AI media narrative in the US that is pointing towards Chinese models as stealing IP or being dangerous or being affiliated with the Chinese government, an authoritarian government.

所有这些都导致了现在对 AI 采取行动的兴趣,但又不知道具体该怎么做。所以可能对所谓的“敌人”采取一种粗糙的手段,我们可以过渡到蒸馏问题。我认为对此有很多讨论。最近 Ben Thompson 终于对蒸馏发表了看法。我认为 Ben 可能是科技领域阅读量最高的博客(Stratechery)。我认为这场辩论,让我们看看从哪里开始。核心问题是:蒸馏有多大帮助,应该怎么做?我一直认为,随着中国模型越来越接近前沿,以及训练机制转向强化学习,蒸馏的影响力正在逐渐减弱。蒸馏通常发生的方式是中国实验室破解 API。“破解”可能是个强烈的词,但他们越狱了 Claude 和 GPT 的 API,以提取推理令牌。

And it’s like all these things are leading up to this moment of interest in taking action on AI and then not really knowing where to do it. So potentially taking a crude instrument to the like quote unquote enemy and we could transition into distillation. I think there’s a lot of discussion on it. Most recently Ben Thompson finally chimed in on distillation. I think Ben is probably one of the is probably the highest read blog in tech (Stratechery). I think that the the debate let’s see where do we even start the debate. The core question is like how much does distillation help and what should you do about it? I’ve been of the opinion that distillation has becoming less and less impactful over time as the Chinese models get closer to the frontier and the trading regime shifts to RL. The way that distillation tends to happen is that the Chinese labs hack the APIs. Hack is like maybe a strong word, but they jailbreak the APIs of Claude and GPT to extract the reasoning tokens.

当你拥有带有工具调用的推理令牌时,那就是完美的 SFT 数据或中期训练数据,可以用来训练基础模型,在重要领域播种一些智能体行为。之后,后训练的核心部分是在智能体领域进行大规模强化学习,以推动前沿发展,这就是他们今天所做的一切。随着 SFT 变得不那么普遍,强化学习只会变得更加普遍。在之前几代中,仅通过扩展 SFT 就能非常接近前沿,如果你能从 Claude 或 GPT 获取一百万次智能体回滚,将其作为你的 SFT 集并进行训练,那将是非常有影响力的。

When you have the reasoning tokens with the tool calls, that is perfect SFT data and/or mid-training data to train the base model with to seed some agentic behaviors in an important domain. And now after that, the core part of post-training is to do large-scale RL in agentic domains to push the frontier and everything that they’re doing today. And RL is only becoming more prevalent with this as SFT becomes less prevalent. In previous generations, you could get very close to the frontier just by scaling up SFT, and that would be really impactful if you could, say, take a million agentic rollouts from Claude or GPT, have that be your SFT set, and train on it.

我认为在过去的几年里,这样做会更能让你接近前沿。本·汤普森说过的话让我非常恼火,他强烈宣称随着强化学习的进行,蒸馏变得越来越有影响力。他在他的文章《谁害怕中国模型?》中这样做了。我们可以在下面链接它。这是一篇公开文章。然后他还在他自己的播客巡演中,他也有播客,说了同样的话。我认为非常重要的是要指出,在强化学习阶段进行蒸馏要困难得多。

I think in previous years that would have done a lot more to get you to the frontier. What Ben Thompson has said, which made me really annoyed, is that he very strongly proclaimed that distillation is getting more impactful as you do RL. He did this in his article 'Who’s Afraid of Chinese Models?' We can link it below. It’s a public one. And then he was also on his own podcast tour. He has a podcast as well, saying the same things. And I think it’s really important to say that distillation during the RL stage is a lot harder.

他所说的是,在强化学习中可以使用的评分模型,本质上你可以让一个模型检查回滚的智能体轨迹,并对不同部分进行评分,看它是否完成了奖励,采取了什么行动。他在暗示中国实验室正在使用 Fable 和 GPT-5.6 等最强模型来在强化学习中进行这种监督。问题在于,大型强化学习运行涉及数百万次回滚。我认为 Thinking Machines 的博客文章提到他们的最终强化学习运行大约有 2000 万到 4000 万次。因此,在像 Fable 或 GPT-5.6 这样的 API 上这样做将极其昂贵,而且可能会成为时间瓶颈,因为这些模型相当慢,坦率地说,可能不会比使用你自己定制的评分模型带来性能提升,还有很多类似的问题。

What he said was that the kind of grading models that can be used during RL, which is essentially you can have a model check over the agentic trajectory of a rollout and grade different parts on if it completed the reward, what actions it took. And he’s insinuating that the Chinese labs are using Fable and GPT-5.6 and the strongest models to actually do this supervision in RL. The problem is that big RL runs are millions and millions of rollouts. I think Thinking Machines’ blog post had like 20 to 40 million or something for their final RL run. So to do this on an API like Fable or GPT-5.6 would be insanely expensive and potentially it would probably be a time bottleneck because these models are pretty slow, and to be frank, might not even give you a performance uplift versus using your own tailored grader model or many things like this.

因此,我认为蒸馏因为强化学习变得更加普遍而更有帮助的论点,并没有基于我们今天所拥有的文献。这对我来说很难,因为本的结论还建议我们应该让禁止蒸馏的服务条款变得非法,我有点想支持他让蒸馏对美国公司合法的激进结论,但我不能支持任何我认为基于错误信息的结论。所以,我也是本的粉丝。如果你是本的粉丝,也能在这件事上推动他,我真的希望你能,因为可能还有一期播客。他会在哪里录制?比如他什么时候录制《Sharp Tech》?周四。

And so I just think the argument that distillation is helping more because RL is becoming more prevalent is not grounded in literature that we have today. This is tough for me because Ben’s article also concludes that we should make terms of service disallowing distillation illegal, which I kind of want to support his radical conclusion to make distillation legal for US companies, but I can’t support any conclusion that I think is based on misguided information. So, I’m also a fan of Ben. If you’re a fan of Ben and could also nudge him on this, I would really appreciate it because there’s probably one more podcast. What is he going to record it on? Like when does he record Sharp Tech? Thursday.

我们必须让他纠正记录,因为我不知道。我觉得最令人恼火的是,科技界最杰出的声音试图成为我们关于蒸馏观点的盟友,即我们应该什么都不做。但这很难。就像他的影响力如此之大,以至于这现在将成为我们必须反驳的现状,我猜这比……我不知道,实际上不,这没有帮助,因为他说蒸馏更重要,这意味着那些对此感到恐惧的人会将其作为数据点,说我们应该采取行动,即使他们不这样做,因为他们可能不会同意他的结论。我不知道。这就是我的咆哮。本,你错了。

We got to get on and get him to correct the record because I don’t know. I find it so annoying that the most prominent voice in tech is trying to be an ally for our point of view on distillation, which is that we should do nothing. But it’s hard. It’s like he has such wide reach that this is now going to be the status quo that we have to debunk, which I guess it’s a better status quo than... I don’t know, actually no, it’s not helpful because he’s saying that distillation is more important, which means the people who are afraid about that are going to use that as a data point to say that we should take action even if they don’t, because they probably won’t agree with his conclusions. I don’t know. That was my rant. Ben, you’re wrong.

00:39:16 Florian Brand: 是的,我认为区分这些阶段很重要。特别是,毫无疑问,它被用于 SFT 阶段,这是后训练的第一个阶段,或者说其中一个阶段。这也是模型学会礼貌举止的地方,比如模型会说“我是 Claude”,就是因为它们在 SFT 阶段学到了这些。这就是个性形成的地方。但强大的能力出现在 RL 阶段,这是花钱的地方,也是你需要一个足够快的评判模型的地方,最好情况下它就在同一批 GPU 上运行,或者非常接近你的 GPU,用一个小型或足够快的模型,这样你就不会因此成为瓶颈。

00:39:16 Florian Brand: Yeah, I do think it is important to differentiate these phases. And especially, there is no doubt that it is used during the SFT stage, which is the first stage of, or one of the stages for, post-training. And that's also where the model picks up its manners, like that's why the models say, "Oh, I am Claude," because they learn this during the SFT stage. That's where this personality is formed. But the strong capabilities come during the RL stage, which is where the money is spent, which is where you need to have a fast enough judge, which in the best case just runs on the same GPUs or very close to your GPUs, with a smallish or fast enough model, so you are not bottlenecked by this.

而且就影响而言,也很难说更好的模型 SFT 数据相对于较弱的模型能带来多大的影响或提升。所以,如果你能从最新的 Claude 模型获得 1000 万 token,而不是两代之前的开源模型,在保持阶段和预训练阶段相同的情况下,这到底能带来多大的提升,这是一个悬而未决的问题,我认为我们不会在论文中看到答案,因为那样你就得展示你的 SFT 和越狱能力。

And in terms of impact, it is also very hard to say how much impact or how much of a boost the better model SFT data gives you versus a lesser model. So if you are able to have 10 million tokens from the latest Claude model versus two generations behind open model, how much of a boost that really gives you, if you keep the stage right the same and the pre-training stage the same, is an open question which I don't think we will see answered in a paper, because then you have to showcase your SFT and your jailbreaking capabilities.

00:40:48 Nathan Lambert: 但我想强调这一点。关于生成 SFT 推理轨迹,已经有相当多的文献。最突出的可能是 OpenAI 的一系列工作,他们做了 Open Thoughts 3 和 Open Thoughts Agent,这些在过去几年里可以说是 Scaling(规模扩张)推理 SFT 的基础性工作。每当有人重新审视这个问题时,他们都没有发现“在你所在领域性能最强的模型是 SFT 的最佳教师”这个答案。人们试过,很多人都试过。这个想法很简单:最先进的开源 SFT 数据集是基于 QwQ-32B 构建的,就像一个古老的推理模型之类的。

00:40:48 Nathan Lambert: But I wanted to double down on this. There's been a good amount of literature on generating SFT reasoning traces. Whether it's the most prominent ones have been Open's line of work, they did Open Thoughts 3 and Open Thoughts Agent, which have kind of been the foundational scaling reasoning SFT works in the last few years. And whenever somebody revisits this question, they have not found the answer that the strongest model on performance in your domain is the best teacher for SFT. People have tried, many people have tried. The idea is so simple: the state-of-the-art open SFT data set is built on QwQ-32B, like an ancient reasoning model or something.

为什么我们不能直接从 GLM 5.2 生成补全,做 SFT,然后改进模型呢?我们不知道。研究就是这样,很多人试过,但这不是一个已解答的研究问题。可能有一些原因,比如基础模型或中期训练与 Qwen 太接近了,所以很难突破。你必须重做中期训练。我认为你必须为推理重做中期训练。我认为推理中期训练和推理 SFT 紧密相连,几乎不需要用不同的词来区分它们。这可能就是问题所在。但文献甚至不知道如何,比如,如果我有一个神奇的 API,能给我来自 Claude/Gemini 的推理轨迹,我实际上不知道在 OLMo 模型上进行微调是否会让 OLMo 更聪明。

Why can we not just generate completions from GLM 5.2, do SFT on it, and improve the model? We don't know. It's like the research, so many people have tried, and it is not an answered research question. There might be something like the base model or the mid-training is too close to Qwen. So therefore it's like hard to break. You have to redo the mid-training. I think you have to redo the mid-training for reasoning. I think reasoning mid-training and reasoning SFT are so closely intertwined. It almost doesn't make sense to have different words for them. That could be the issue. But the literature doesn't even know how to, like, if I had a magical API that gave me reasoning traces from Claude/Gemini, I actually don't know if fine-tuning an OLMo model on that would make OLMo smarter.

这是最疯狂且未解答的研究问题之一。而这恰恰让蒸馏这件事变得很有趣,就像,是的,我认为中国实验室确实使用像 Opus 这样的强模型来生成一些 SFT 数据,但他们也在创新。我希望他们能告诉我们如何让这该死的东西起作用。我认为这种范式是:OpenAI 和 Anthropic 找到一个他们做得非常好的细分领域,然后中国实验室可以在那里获取一些样本,以启动他们的数据引擎,这样你就能在特定领域获得几个月的领先优势。

It's one of the most wild unanswered research questions. And this just makes the distillation thing so funny, where it's like, yes, the Chinese labs, I think, are using strong models like Opus for some SFT data, but they're also innovating. I would love for them to tell us how to make this freaking work. And I think the paradigm is like OpenAI and Anthropic find a niche domain that they do so well at, and then the Chinese labs can get some samples there to kind of bootstrap their data engine, and that's where you will gain a few months on a specific domain.

但在数学、代码和 Terminal-Bench 等核心领域进行爬山式改进,他们只是在做同样的事情,这非常困难:生成那些对当前模型来说具有挑战性、并提供真实且无奖励黑客行为的学习环境的问题。这就是前沿数据研究目前的状况。生成这些难题确实很难。我相信中国的研究机构也在做同样的事情。我不知道。这就是我的抱怨。我有点忘记我们谈话的上下文了。

But hill-climbing on these core domains like math and code and Terminal-Bench, they're just doing the same thing, which is so hard: generating prompts that are problems with environments that are hard for the current models and provide real, non-reward-hacking learning behavior. And that is what frontier data research looks like right now. And it is hard to generate these hard problems. And I'm sure the Chinese labs are doing the same things. And I don't know. That's my rant. I've kind of lost the context of our conversation.

00:43:23 Florian Brand: 不,不,我同意。或者总结一下,是的,SFT 或蒸馏有一定效果。是的,它给了它们提升,但没有人们希望或似乎认为的那么多。

00:43:23 Florian Brand: No, no, I would agree. Or to recap, yeah, SFT or distillation has some effect. Yeah, it gives them a boost, but not that much as people would like or seem to think it gives.

00:43:39 Nathan Lambert: 我认为这是对对话的一个很好的总结。这也变得有点令人厌烦,因为有人说所有这些开放模型之所以好,只是因为它们在蒸馏,这绝对不是事实,因为如果是这样,每个人都能轻易地通过使用 GLM 或 K3 的数据进行蒸馏来赶上它们。但我们没有看到,或者我们不会仅从 SFT 中看到这一点。

00:43:39 Nathan Lambert: I think that's a good summary of the conversation. It also becomes kind of tiresome because it says that all these open models are just good because they are distilling, which definitely isn't the case, because if it were the case, everyone would easily be able to catch up to a GLM or to a K3 by using its data for distillation. But we have not, or we won't see this from SFT alone.

00:44:12 Florian Brand: 是的,我同意。你有什么预测或想讨论的其他话题吗?

00:44:12 Florian Brand: Yeah, I agree. Do you have any predictions or more topics you want to get to?

00:44:18 Nathan Lambert: 在预测方面,我认为我们,或者我重新审视了我们去年的预测,基本上是说一切都会像前一年一样继续。我们预测会看到更大的模型,超过两万亿参数,这确实发生了,但我认为今年模型规模不会有更大的爆炸。我们可能会看到总参数略超三万亿的模型,但我预计不会有五万亿或十万亿参数的模型,而且今年我们开源这样的模型会让我非常惊讶。然后去年的列表,我们能重新做吗?我们不必做全部。

00:44:18 Nathan Lambert: In terms of predictions, I think we are, or I revisited ours from last year, and it basically said everything will continue like it did the previous year. We predicted that we will see bigger models up over two trillion parameters, which it did, and I don't think we will see a much bigger explosion in terms of model size this year. We might see some model a bit bigger than three trillion parameters total, but I don't expect a five or ten trillion parameter model, and we open this year that would really surprise me. Then list from last year, can we redo this? We don't have to do the whole thing.

00:45:03 Nathan Lambert: 这就是我们在 2025 年底的情况。现在你认为谁属于前沿模型?嗯,是 Kimi 和智谱。DeepSeek 这些天有点难啃。我觉得他们应该算作紧密竞争者。所以我会把 DeepSeek 和通义千问列为紧密竞争者,而 Kimi 和智谱属于前沿。你觉得还有谁配得上紧密竞争者吗?因为在那之后,值得关注的和更低的层级,有太多了。

00:45:03 Nathan Lambert: This is where we were at the end of 2025. Who do you put in frontier now? Well, it is Kimi and it is Zhipu. DeepSeek is kind of a hard nut these days. Like, I think they would be in close competitors. So I would put DeepSeek and Qwen as close competitors, with Kimi and Zhipu as frontier. Do you think anyone else would deserve close competitor? Because after that, noteworthy and below, there are so many.

00:45:38 Florian Brand: 我认为到年底我们会看到 MiniMax 带来惊喜。我认为我们会看到一个大型模型,一个真正的大模型,不是 M3 规模,而是万亿参数以上,这将在输出方面让我们对 MiniMax 刮目相看。所以到年底我仍然会把他们列为紧密竞争者。

00:45:38 Florian Brand: I think we will see a surprise from MiniMax by end of the year. I think we will see a big model, like a really big model, not M3 size but trillion parameters plus, which will surprise us in terms of the outputs of MiniMax compared to before. So I would still put them at close competitors by the end of the year.

00:45:54 Nathan Lambert: 你认为到年底会有美国公司进入紧密竞争者吗?Nemotron?我不认为我会把它放进去。Thinking Machines 更接近,特别是如果较小的模型真的取得突破的话。但我认为现在还不能把他们放进去。据说 Reflection 只有在拥有前沿模型时才会发布。但问题是,我们能等到吗?比如,我们认为到年底会有美国公司进入这个大致相当于我们前五的名单吗?所以前五名还是那些,但顺序会打乱。

00:45:54 Nathan Lambert: Do you think any US companies will be in the closed competitors by end of the year? Nemotron? I don't think I would put there. Thinking Machines closer, especially if the smaller model really breaks through. But I don't think I would put them there yet. Reflection is supposedly only wants to release if they have a model that's frontier. But then the question is, will we get it? Like, do we think that any US companies will get into this, what is roughly like our top five by the end of the year? So the top five are the same but reshuffled.

00:46:36 Florian Brand: 我认为有可能他们真的非常接近。这也取决于我们认为什么才算接近。比如,我认为 Nemotron 和 Thinking Machines 会发布一些模型,它们作为很好的基础模型,可以针对你的领域进行微调,这并不意味着它们能像前沿模型那样直接使用,但它们有很高的实用性,所以我会把它们列为紧密竞争者,因为你只需要找到你的数据,然后把模型推向正确的方向。

00:46:36 Florian Brand: I would say it is possible that they are really close. It also depends on what we think matters for closeness. Like, I think Nemotron and Thinking Machines will release models which act as really good base to be fine-tuned for your domain, which doesn't mean they are usable like a frontier model, but they have so much utility that I would put them into close competitors because you would just need to find your data and push the model into the right direction.

00:47:12 Nathan Lambert: 我本来想,到那时我们会把这个名单扩成六家,包括一家美国公司。比如,如果我们在 11 月底做这个,我猜会有一家美国公司,最可能是 Nvidia、Thinking Machines 或 Reflection,做出一些事情,让我们认为有一家美国公司在这个顶级集群里,这将是相当长一段时间以来的第一次。

00:47:12 Nathan Lambert: As I was going to think that we would make this a group of six with a US company by then. Like, if we do this in late November, I would guess that a US company, most likely Nvidia, Thinking Machines, or Reflection, does stuff that gets us to say that there is an American company in this top cluster, which would be a first time for a while.

00:47:43 Florian Brand: 是的,我认为这是现实的。我的一张王牌是腾讯,我认为我们可能在年底前看到一些东西。他们有了新的领导层。这次他们在 Apache 许可下发布了混元模型。等等,所以腾讯一直有这些自定义许可证,禁止英国、韩国以及整个欧盟的任何人使用他们的模型,并且还有使用政策等。而随着混元和新领导层的到来,他们拥有了一个约 2500 亿参数的真正有能力的模型。我认为到年底我们可能会看到一个大型模型的发布,这将让那些不关注生态系统的人感到惊讶。

00:47:43 Florian Brand: Yeah, I think that is realistic. My one wild card is Tencent, which I think we might see something by end of the year. They got some new leadership. They released their Hunyuan model under Apache this time. Wait, so Tencent always had these custom licenses which disallowed anyone in the UK and South Korea and the entirety of the EU to use their model and also had acceptance use policy and so on. And with Hunyuan and their new leadership, they got a really competent model at 250ish billion parameters. And I think by end of the year we might see a big model release which will surprise the people not following the ecosystem.

00:48:36 Nathan Lambert: 是的,我也确信我们会遇到一些惊喜。这总是 AI 的事情,尤其是开放模型。这非常非常不可预测。好的,我认为这是一个很好的停止点。我们可能真的应该每季度做一次。这并不难,人们会喜欢的。但是很高兴见到你,我们很快再聊。希望很快能面对面。

00:48:36 Nathan Lambert: Yeah, I am also sure we will be in for some surprises. This is always the thing with AI and especially open models. It’s very very unpredictable. Okay, I think this is a good place to stop. We probably should really do this quarterly. It’s not that hard and people will enjoy it. But good to see you and we’ll talk soon. Hopefully in person soon.

[](https://substack.com/profile/427885987-subrata)[](https://substack.com/profile/4843531-jake-flomenberg)[](https://substack.com/profile/400785837-wa)[](https://substack.com/profile/19712584-v-p)[](https://substack.com/profile/3346222-henry-ogedegbe-jr)

[](https://substack.com/profile/427885987-subrata)[](https://substack.com/profile/4843531-jake-flomenberg)[](https://substack.com/profile/400785837-wa)[](https://substack.com/profile/19712584-v-p)[](https://substack.com/profile/3346222-henry-ogedegbe-jr)

关于本集的讨论 Discussion about this episode

我很好奇你对更小模型的看法,比如 Liquid AI 的模型?

I am curious what is your take on even smaller models, like ones from Liquid AI?

它们还行,但不会对 AI 的发展轨迹产生太大影响。前沿实验室如果想做,可以做出更好的小模型。

They're fine, don't really impact the trajectory of AI much. The frontier labs can make much better small models if they want to.

关于 AI 最新发展的音频文章,以及对领域内顶尖科学家的访谈。打破炒作,理解内在原理,讲述故事。

Audio essays about the latest developments in AI and interviews with leading scientists in the field. Breaking the hype, understanding what's under the hood, and telling stories.

关于 AI 最新发展的音频文章,以及对领域内顶尖科学家的访谈。打破炒作,理解内在机制,讲述故事。

Audio essays about the latest developments in AI and interviews with leading scientists in the field. Breaking the hype, understanding what's under the hood, and telling stories.

[](https://substack.com/@natolambert?utm_source=author-byline-face-podcast)

[](https://substack.com/@natolambert?utm_source=author-byline-face-podcast)

[](https://substack.com/@xeophon?utm_source=author-byline-face-podcast)

[](https://substack.com/@xeophon?utm_source=author-byline-face-podcast)

2025 年 12 月 18 日 • Nathan Lambert 和 Florian Brand

Dec 18, 2025 • Nathan Lambert and Florian Brand

互动版:图/公式 + 针对本篇提问 →