本文分析了智谱 AI 发布的 GLM-5.3 模型,该模型在智能体编码基准上达到前沿性能,仅用约 750B 参数,仅为 Kimi K3 等竞争对手的三分之一。作者认为,中国实验室与美国同行保持同步并非主要依靠蒸馏,而是通过更快的发布周期、战略性的基准聚焦和卓越的后训练技术。关键因素包括智谱 AI 能在数天内而非数月内发布模型,从而持续进行基准爬山,并针对高价值用例进行重点优化。文章还强调了中国日益增长的强化学习数据产业和智谱 AI 的计算效率。作者总结道,尽管实施了请求分类器等安全措施,但随着模型规模缩小和开放权重更易获取,强大网络能力的扩散不可避免,并敦促政府或联盟提供产业级指导,为这一转变做好准备。
This article analyzes the release of Z.ai's GLM-5.3 model, which achieves frontier-level performance on agentic coding benchmarks with only ~750B parameters, a third of competitors like Kimi K3. The author argues that Chinese labs keep pace with American counterparts not primarily through distillation, but through faster release cycles, strategic benchmark focus, and exceptional post-training expertise. Key factors include Z.ai's ability to release models in days rather than months, allowing continuous benchmark hill-climbing, and their targeted focus on high-value use cases. The article also highlights the growing RL data industry in China and Z.ai's compute efficiency. The author concludes that while safety measures like request classifiers are implemented, the proliferation of strong cyber capabilities is inevitable as model sizes shrink and open weights become more accessible, urging industrial-scale government or coalition guidance to prepare for this transition.
核心贡献 · Key contributions
GLM-5.3 以约 7500 亿参数达到前沿智能体式编码性能,仅为 Kimi K3 等竞争对手的三分之一。 GLM-5.3 achieves frontier-level agentic coding performance with ~750B parameters, a third of competitors like Kimi K3.
中国实验室通过更快的发布周期(数天对比数月)保持同步,实现持续的基准测试爬山。 Chinese labs keep pace via faster release cycles (days vs months), enabling continuous benchmark hill-climbing.
Z.ai 的优势在于后训练,利用更多环境、多样化任务和算力进行强化学习。 Z.ai's strength lies in post-training, using more environments, diverse tasks, and compute for RL.
战略性基准测试聚焦和高价值用例定位使较小模型能有效竞争。 Strategic benchmark focus and targeting high-value use cases allow smaller models to compete effectively.
中国快速发展的强化学习数据产业(部分由美国数据公司推动)支持了中国实验室。 The growing RL data industry in China, partly driven by American data companies, supports Chinese labs.
Z.ai 算力效率高,并受益于与清华大学人才库的紧密联系。 Z.ai is compute-efficient and benefits from close ties to Tsinghua University's talent pool.
局限 · Limitations
GLM-5.3 可能比美国前沿模型更窄,专注于智能体式编码和纯文本。 GLM-5.3 is likely narrower than American frontier models, focusing on agentic coding and text-only.
请求分类器和思维链监控等安全措施对开放权重可能不足。 Safety measures like request classifiers and chain-of-thought monitoring may be insufficient for open weights.
文章推测蒸馏但缺乏直接证据,能力对等仍未解释。 The article speculates on distillation but lacks direct evidence, leaving capability parity unexplained.
由于潜在的微妙基准测试优化,基准分数可能无法完全反映真实世界性能。 Benchmark scores may not fully reflect real-world performance due to potential subtle benchmaxxing.
随着模型规模缩小和开放权重传播,强大网络能力的扩散不可避免。 The proliferation of strong cyber capabilities is inevitable as model sizes shrink and open weights spread.
论文章节 · Sections(共 1)
提示:这其实不是一个蒸馏的故事。Hint: It’s really not a distillation story.