本文讨论了开放权重 AI 模型的快速进展,重点关注了 Kimi K3 和 Qwen 3.8 等最新发布,及其对 AI 格局的影响。作者 Nathan Lambert 和 Florian Brand 分析了开放与封闭模型之间的性能差距,认为基准测试常常因评估方法不同而误导对这一差距的认知。他们强调了中国模型的惊人质量,并将其归因于资本效率以及在计算、数据和人才方面的战略投资。讨论涵盖了大型模型后训练的挑战、蒸馏在模型开发中的作用,以及推动开源采用的地缘政治和经济因素。作者总结道,尽管开放模型在编程等特定任务上正在缩小差距,但在长尾能力上仍显不足,且生态系统正迅速专业化以支持这些更大的模型。他们预测开放模型的发布将继续加速,并强调细致评估优于简单比较的重要性。
This article discusses the rapid advancements in open-weight AI models, focusing on recent releases like Kimi K3 and Qwen 3.8, and their implications for the AI landscape. The authors, Nathan Lambert and Florian Brand, analyze the performance gap between open and closed models, arguing that benchmarks often misrepresent this gap due to varying evaluation methods. They highlight the surprising quality of Chinese models, attributing it to capital efficiency and strategic investments in compute, data, and talent. The discussion covers the challenges of post-training large models, the role of distillation in model development, and the geopolitical and economic factors driving open-source adoption. The authors conclude that while open models are closing the gap in specific tasks like coding, they still lag in long-tail capabilities, and the ecosystem is rapidly professionalizing to support these larger models. They predict continued acceleration in open model releases and emphasize the importance of nuanced evaluation over simplistic comparisons.
核心贡献 · Key contributions
分析开源与闭源模型的差距,认为开源模型在关键基准上仅落后数月。 Analyzes the open-closed model gap, arguing open models are only months behind on key benchmarks.
讨论中国实验室的资本效率和算力获取,解释其快速进步的原因。 Discusses Chinese labs' capital efficiency and compute access, explaining their rapid progress.
驳斥蒸馏神话,指出与强化学习创新相比,SFT 蒸馏的影响有限。 Debunks distillation myths, stating SFT distillation has limited impact compared to RL innovations.
强调禁止开源模型的网络安全风险,引用防御者依赖开源权重的情况。 Highlights cybersecurity risks of banning open models, citing defenders relying on open weights.
预测前沿模型梯队,Kimi 和智谱领先,年底可能出现美国入局者。 Predicts frontier tier list with Kimi and Zhipu leading, and potential US entrants by year-end.
局限 · Limitations
基准比较可能无法反映真实世界任务性能的差异。 Benchmark comparisons may not capture real-world task performance differences.
关于中国实验室效率的说法缺乏算力和成本的公开数据。 Claims about Chinese labs' efficiency lack public data on compute and costs.
蒸馏辩论依赖轶事证据,而非同行评审研究。 Distillation debate relies on anecdotal evidence, not peer-reviewed studies.
预测具有推测性,可能因 AI 快速发展而失效。 Predictions are speculative and may be invalidated by rapid AI developments.
关注编码和智能体任务可能忽视其他重要模型能力。 Focus on coding and agentic tasks may overlook other important model capabilities.
论文章节 · Sections(共 3)
开放模型回顾:更多关于 Kimi K3、Qwen 3.8、习近平 WAIC 演讲、蒸馏、开放与封闭差距以及未来展望Open models recap: more on Kimi K3, Qwen 3.8, Xi's WAIC speech, distillation, the open-closed gap, and what's next