GLM-5.3: How Chinese labs keep stride with the frontier
打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→本文分析了智谱 AI 发布的 GLM-5.3 模型,该模型在智能体编码基准上达到前沿性能,仅用约 750B 参数,仅为 Kimi K3 等竞争对手的三分之一。作者认为,中国实验室与美国同行保持同步并非主要依靠蒸馏,而是通过更快的发布周期、战略性的基准聚焦和卓越的后训练技术。关键因素包括智谱 AI 能在数天内而非数月内发布模型,从而持续进行基准爬山,并针对高价值用例进行重点优化。文章还强调了中国日益增长的强化学习数据产业和智谱 AI 的计算效率。作者总结道,尽管实施了请求分类器等安全措施,但随着模型规模缩小和开放权重更易获取,强大网络能力的扩散不可避免,并敦促政府或联盟提供产业级指导,为这一转变做好准备。
This article analyzes the release of Z.ai's GLM-5.3 model, which achieves frontier-level performance on agentic coding benchmarks with only ~750B parameters, a third of competitors like Kimi K3. The author argues that Chinese labs keep pace with American counterparts not primarily through distillation, but through faster release cycles, strategic benchmark focus, and exceptional post-training expertise. Key factors include Z.ai's ability to release models in days rather than months, allowing continuous benchmark hill-climbing, and their targeted focus on high-value use cases. The article also highlights the growing RL data industry in China and Z.ai's compute efficiency. The author concludes that while safety measures like request classifiers are implemented, the proliferation of strong cyber capabilities is inevitable as model sizes shrink and open weights become more accessible, urging industrial-scale government or coalition guidance to prepare for this transition.
杂务说明:我正在旅行,因此无法为这篇文章录制配音。编辑——我在发出邮件后添加了关于中国数据行业的第 5 点。
Housekeeping: I'm traveling so cannot make a voiceover for this post. EDIT — I added a bullet point 5 on the Chinese data industry after sending the email out.
今天,Z.ai 发布了他们的 GLM-5.3 模型,目前仅在编程计划中可用,即将在其 API 中推出,并将在两周后上线 Hugging Face(开放权重)。该模型表现卓越,分数提升令人震惊。在许多基准测试中,该模型已超越 Moonshot AI 的 Kimi K3,并在某些基准上超越了 Claude Fable 5 或 GPT-5.6-Sol。
Today, Z.ai announced their GLM-5.3 model, currently only available in the coding plan, coming soon to their API and in two weeks' time to Hugging Face (open weights). This model looks exceptional, with a somewhat astounding increase in scores. On many benchmarks the model has surpassed Moonshot AI's Kimi K3 and on some it's surpassed Claude Fable 5 or GPT-5.6-Sol.
这使得该模型大致处于智能体式编码基准的前沿,而参数仅约 750B——只有 Kimi K3 的三分之一!Z.ai 的博客文章相当直白,开头便是一句大胆的话:
This puts the model more or less at the frontier of agentic coding benchmarks, with only ~750B parameters – a third of Kimi K3! The Z.ai blog post is rather straightforward, and starts with a bold sentence:
GLM-5.3 与 GLM-5.2 使用相同的基础模型,但后训练大幅扩展。冒一点过度简化的风险,Z.ai 在后训练方面似乎有优势,而 Kimi 更像是预训练的杰作。此次发布后,有很多讨论在问:中国怎么能跟得这么好?这么小的模型怎么能与领先的美国公开模型匹敌?这些结果是真的吗?
GLM-5.3 is the same base model as GLM-5.2 with substantially extended post-training. To risk a broad oversimplification, Z.ai seems to have a strength in post-training when compared to Kimi, which is more of a pretraining masterpiece. Following this release there have been a lot of discussions wondering how China can keep up so well? How can such a small model be matching the leading public American models? Are these results real?
最简单的解释是,Z.ai 非常擅长他们所做之事——值得回顾的是,他们研究这一系列模型的时间比业内几乎任何人都长。以下是 GLM 模型的简要历史。
The simplest explanation is that Z.ai is very good at what they do – it’s worth recalling that they’ve been working on this line of models longer than almost anyone in the industry. Here’s a brief history of the GLM models.
* GLM(通用语言模型)——2021 年 3 月——由清华大学数据挖掘/知识工程组(THUDM)发布。权重 * GLM-130B——2022 年 8 月——扩展版本。GLM-130B 至 GLM-4 的技术报告——权重 * ChatGLM——2023 年 3 月 14 日——首个聊天版本。权重
* GLM (General Language Model) — March 2021 — released by THUDM, Tsinghua University’s Data Mining / Knowledge Engineering group. Weights * GLM-130B — August 2022 — Scaled version. Technical report for GLM-130B through GLM-4 — Weights * ChatGLM — March 14, 2023 — first chat version. Weights
GLM-5.3 与 GLM-5.2 使用相同的基础模型,但后训练大幅扩展。冒着重度简化的风险,Z.ai 在后训练方面似乎比 Kimi 更有优势,而 Kimi 更像是预训练的杰作。此次发布后,引发了许多讨论:中国如何能保持如此领先?如此小的模型怎能媲美美国领先的公开模型?这些结果是真的吗?
GLM-5.3 is the same base model as GLM-5.2 with substantially extended post-training. To risk a broad oversimplification, Z.ai seems to have a strength in post-training when compared to Kimi, which is more of a pretraining masterpiece. Following this release there have been a lot of discussions wondering how China can keep up so well? How can such a small model be matching the leading public American models? Are these results real?
最简单的解释是 Z.ai 非常擅长他们所做之事——值得回顾的是,他们研发这一系列模型的时间比业内几乎任何人都长。以下是 GLM 模型的简要历史。
The simplest explanation is that Z.ai is very good at what they do – it’s worth recalling that they’ve been working on this line of models longer than almost anyone in the industry. Here’s a brief history of the GLM models.
* ChatGLM3 — 2023 年 10 月 27 日 — 权重
* ChatGLM3 — October 27, 2023 — Weights
* GLM-4 — 2024 年 1 月 16 日 — 更名为 GLM;随后于 6 月发布开放权重的 GLM-4-9B。权重
* GLM-4 — January 16, 2024 — rebranded as just GLM; open-weight GLM-4-9B followed in June. Weights
* GLM-5 — 2026 年 2 月 11 日 — 最新主要版本。权重
* GLM-5 — February 11, 2026 — latest major generation. Weights
今年 6 月 22 日发布的 GLM 5.2 引起了广泛关注——发布数周后,我经常听到我认识的 AI 研究人员仍在用这个模型,因为它速度快(有些人在内部集群上部署该模型以获得比公开服务更快的速度)且简单(作为一个没有回滚等功能的模型,在处理前沿 AI 系统时)。GLM-5.2 完全不负众望。
GLM 5.2, released on June 22 of this year, was a big deal – weeks after the release, I regularly heard from AI researchers I know who still used the model due to its speed (some deploy the model on internal clusters for faster speeds than public offerings) and simplicity (as a model with no rollbacks, etc., when working on frontier AI systems). GLM-5.2 altogether stood up to the hype.
我自己也经历过类似的否认,心想:“他们是怎么做到的?_当然_这些模型不可能像看起来那么好。”美国公司拥有如此巨大的资源领先优势,但在能力上却似乎无法拉开差距,这有点令人不安。常见的解释是蒸馏,我对此写过很多,但我认为这不是主要因素。关于这一点,最近有一篇论文展示了从前沿模型中提取推理轨迹的简单方法——这正是中国实验室肯定可以大规模使用的东西。我不明白为什么美国实验室没有更快地修补这种行为,反而跑去向政府寻求政策帮助。这对我来说说不通。
I’ve been going through some of the same denial myself, thinking “how do they keep doing this? _Surely_ the models aren’t as good as they look.” There’s something a bit off-putting with how the American companies have such a commanding resource lead, but can’t seem to pull away in capabilities. The common answer is distillation, which I’ve written at length about, but I deem not to be the major factor. On that note, there was a recent paper that showed simple methods for extracting the reasoning traces from frontier models – this is the sort of thing that Chinese labs could definitely use at scale. I’m confused why the labs in the U.S. haven’t patched this behavior faster; instead they’re running to the government asking for policy help. It doesn’t add up for me.
Z.ai 的博客直截了当,与以强化学习为主导的训练机制相符。他们表示使用了“更多的环境、更多样化的任务,以及在这些任务上投入更多的算力进行训练”。人们不能简单地“蒸馏”强化学习环境、大规模运行这些环境的基础设施,或有效混合它们的算法。
Z.ai's blog is direct and matches with an RL-dominated training regime. They say they used "more environments, more diverse tasks, and more compute spent training on them." One does not simply "distill" RL environments, infrastructure to run them at scale, or algorithms to mix them together effectively.
Interconnects AI 是一家由读者支持的出版物。请考虑成为订阅者。
Interconnects AI is a reader-supported publication. Consider becoming a subscriber.
那么,如果不是蒸馏,中国实验室是如何做到的呢?他们是在“刷榜”(benchmaxxing)吗?公认的“刷榜”定义是让模型专注于测试集,使得真实世界性能与纸面分数存在显著差异。决定因素更多是宏观层面的,而非技术层面(是的,技术细节确实重要,但很难区分不同实验室之间的差异):
So, how do the Chinese labs do it if not distillation? Are they benchmaxxing? An accepted definition of benchmaxxing is focusing the model on the test sets, such that the real-world performance meaningfully differs from the on-paper scores. The determining factors are much more big picture than technical (yes, the technical details definitely matter, but are harder to differentiate from lab to lab):
1. Z.ai 的发布周期可能只有几天,而不是像 OpenAI 或 Anthropic 那样需要几个月。极有可能的是,OpenAI 和 Anthropic 的内部模型远优于 Z.ai 和月之暗面(Moonshot AI)。尽管如此,这些美国公司往往需要数月时间才能向公众发布模型,这在前沿模型的采用决策中极大地有利于中国实验室。简而言之——中国实验室利用美国实验室在发布前测试的所有时间来持续在基准测试上“爬山”(SpaceXAI 在这方面可能更接近中国实验室)。鉴于进展速度如此之快,这很可能是中国实验室保持前沿地位的最大决定因素。到目前为止,这对美国实验室来说在经济上是可以接受的,因为他们的模型仍有大量需求。
1. The time to release for Z.ai is likely days, not months as with OpenAI or Anthropic. It is very, very likely that OpenAI and Anthropic have far better internal models than Z.ai and Moonshot AI. Still, these American companies tend to take months to release their models to the public, which massively flatters the Chinese labs in adoption decisions at the frontier. To put it simply – the Chinese labs use all the time that American labs do pre-release testing to keep hillclimbing on benchmarks (SpaceXAI is likely far closer to the Chinese labs here). With the pace of progress being so fast, this is likely the largest determining factor of why Chinese labs stay at the frontier. This, so far, has been economically acceptable for the American labs, as they've still had massive demand for their models.
随着构建大语言模型的实验室内部模型自我改进循环的加速,如果这些反馈循环中的任何一个需要用户数据,这种更快的发布周期可能会极大地有利于中国实验室,使他们的产品在下一个更优越的模型出现之前拥有更长的生命周期,从而削弱对其模型的需求。
As model self-improvement loops ramp up within the labs building LLMs, if any of these feedback loops require user data, this faster release cycle could massively favor the Chinese labs, giving their offerings longer lifespans before the next vastly superior model comes out, undercutting demand for their models.
这些显然就是业内许多人担心的竞赛动态。在如此多的实验室都在领先能力范围内构建前沿模型的情况下,很难看到这种情况在不久的将来会有所缓和。
These are very clearly the race dynamics that many in the industry worry about. With so many labs building frontier models in the envelope of leading capabilities, it is hard to see this abating in the near future.
2. 是的,Z.ai 可能比 OpenAI 或 Anthropic 更关注公开基准测试。这些基准测试,例如在人工智能分析智能指数或类似聚合器上获得高分,会直接影响其股价。他们在很多方面都需要这样做,以持续筹集资金并保持团队士气,因为作为挑战美国巨头的后起之秀,这是一个绝佳的故事。
2. Yes, Z.ai probably cares slightly more about public benchmarks than OpenAI or Anthropic. These benchmarks, e.g. scoring highly on the Artificial Analysis Intelligence Index, or similar aggregators, have a very direct impact on their stock price. They in many ways need to do this to keep raising capital and maintain team morale, as being the scrappy underdog matching American giants is a wonderful story.
微妙的基准测试优化(benchmaxxing)并不一定源于绝望或类似的压力。这在众多实验室中已成为行业标准。许多公司的数据采集策略是购买他们在基准测试中落后的数据。
Subtle benchmaxxing does not need to come out of desperation or any similar pressures. It’s the industry standard across a remarkable number of labs. Many companies’ data acquisition strategy is to buy data on the benchmarks they’re behind on.
3. Z.ai 并没有将基准测试优化到 GLM-5.3 被“烤焦”的程度(至少不是故意的,而且他们会检查这一点)。目前每个实验室都在应对扩展强化学习(RL)的棘手问题。Anthropic 的 Opus 5 和 Sonnet 5 模型尽管取得了惊人的基准分数,但声誉却褒贬不一。业内所有人都面临同样的困境,因此有些模型权重最终比其他模型更容易使用,但他们发布博客中的基准分数是真实可靠的。
3. Z.ai is not benchmaxxing to the point where GLM-5.3 is fried (at least not intentionally, and they’ll check for it). Every lab is dealing with the rough edges of scaling RL right now. Anthropic’s Opus 5 and Sonnet 5 models have very mixed reputations, despite the incredible benchmark scores. Everyone in the industry is in the same boat, so some model weights end up being easier to use than others, but the benchmark scores in their release blogs are the real deal.
4. GLM-5.3 可能比 Claude Fable 或 GPT Sol 更窄。当 GPT-5.2 发布时,除了智能体式编码之外,评价褒贬不一。与此同时,OpenAI 和 Anthropic 支持着拥有无数用例的大型企业。这是处于采用曲线早期阶段的公司的一个优势——你可以瞄准最有价值的用例。在后训练中,少关注一点会让最终模型的组装变得容易得多。
4. GLM-5.3 is likely a narrower model than Claude Fable or GPT Sol. When GPT-5.2 was released, it had mixed reviews outside of agentic coding. At the same time, OpenAI and Anthropic support very large businesses with countless use-cases for their models. This is a benefit of being a company earlier in their adoption curve – you can target the most valuable use-cases. Within post-training, caring about a bit less will make assembling the final model _far_ easier.
我有点夸大其词,因为据报道,Z.ai 凭借强大的本地部署业务,年经常性收入(ARR)已达到 10 亿美元。
I’m overstating this a bit, as Z.ai reportedly reached $1B of ARR on the back of a strong on-premises deployment business.
同样,旗舰 GLM 模型目前不具备视觉能力。纯文本模式确实有助于 Z.ai 获得更具竞争力的分数,但这也是一个竞争更激烈的领域。另一方面,像 Inkling-Small 这样的模型则被设计为全模态模型。
Similarly, the flagship GLM models have not had visual capabilities. Being text-only definitely helps Z.ai get more competitive scores, but it is a more competitive space. On the other side of things are models like Inkling-Small, which is designed to be omnimodal.
5. (新增)强化学习数据产业正在中国兴起。我们关注的许多消息来源和传闻都提到,中国的数据产业正在蓬勃发展——很大程度上是由美国数据公司向中国模型实验室销售所推动的。这可能表现为中国实验室购买许多与美国前沿实验室相同的强化学习环境,并更早地发布下游经过强化学习的模型。我们对该市场的规模和影响仍存在很大的不确定性,但它无疑正变得重要。
5. (ADDED) The RL data industry is taking off in China. Many sources and rumor-mills we’re following have been mentioning how the data industry is taking off in China — very much driven by American data companies selling to Chinese model labs. This could look like Chinese labs buying many of the same RL environments that are used by American frontier labs, and releasing the downstream RL’d model sooner. We still have large error bars on the scale and impact of this market, but it is certainly becoming important.
6. Z.ai 是一家技术极为精湛的大语言模型(LLM)组织——其算力效率可能远超 OpenAI 和 Anthropic。这一点需要反复强调。这些人非常擅长自己的工作。该公司与清华大学关系密切,而清华大学拥有许多中国最优秀的计算机科学家。这个丰富且渴望学习的人才库,与任何西方同行一样,是他们成功的关键。
6. Z.ai is an extremely skilled LLM organization – one that is likely far more compute efficient than OpenAI / Anthropic. This needs repeating. These folks are very good at what they do. The company has very close ties to Tsinghua University, which is home to many of the best Chinese computer scientists. This abundant, eager talent pool is as central to their success as it is for any Western counterpart.
总而言之,他们用 GLM 系列模型执行的似乎是一个非常完美的策略。恭喜发布!我很期待权重公开,这样我就能进行更深入的测试(我倾向于使用美国的开放权重推理服务,如 Fireworks 或 Baseten)。
Altogether, it seems like a perfectly good strategy they’re executing with the GLM line of models. Congrats on the release! I’m excited for the weights to be out so I can do more extended testing (I tend to use American open-weight inference services like Fireworks or Baseten).
这是朝着不可避免的、强大的网络能力在经济中扩散的又一步。Z.ai 已经承认了这一点,并表示:
This is another step towards the inevitable proliferation of very strong cyber capabilities across the economy. Z.ai has acknowledged this, saying:
他们接着承认,他们如何通过请求分类器和思维链监控(在模型对齐之上)来监控其平台上的推理。这里的细节决定成败,目前尚不清楚每个 AI 实验室的执行水平如何。能力的扩散取决于最低的共同标准。
They go on to acknowledge how they’re monitoring inference on their platforms via a request classifier and chain of thought monitoring (on top of model alignment). The devil is in the details here, and it is unclear the level of execution every AI lab will have here. The capability diffusion is determined by the lowest common denominator.
归根结底,当真正的开放权重模型即将到来时,这类安全措施几乎无关紧要。如果不是 GLM-5.3,也会有其他模型。拥有这些能力的模型规模正在随时间缩小,变得更容易修改和部署(可能没有安全保障)。Z.ai 做了一些正确的事情,包括推动更多的漏洞发现和主动管理,但任何单一公司都远不能独自应对这一问题。
At the end of the day, this type of safety barely matters when true open-weights are coming. If not GLM-5.3, then another model. The size of the models with these capabilities is reducing over time, becoming easier to modify and deploy (potentially without safeguards). Z.ai does some of the right things, including pushing for more vulnerability discovery and proactive management, but any single company is far from being able to handle this on their own.
我们需要由政府或行业联盟主导的工业级指导,立即为这一转变在所有软件中做好准备。
We need industrial-scale guidance led by the government or industry coalitions to immediately prepare for this transition across all software.