Kimi K2:开放智能体智能

Kimi K2: Open Agentic Intelligence

杨植麟 Zhilin Yang · Moonshot AI · 2025-07-28 · arXiv:2507.20534 ↗ · 被引 316

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

我们推出了 Kimi K2,一个混合专家(MoE)大语言模型,具有 320 亿激活参数和 1 万亿总参数。我们提出了 MuonClip 优化器,通过新颖的 QK-clip 技术改进了 Muon,解决了训练不稳定性问题,同时保留了 Muon 的高级 token 效率。基于 MuonClip,K2 在 15.5 万亿 token 上进行了预训练,且无损失尖峰。在后训练阶段,K2 经历了多阶段后训练过程,重点是大规模智能体数据合成流水线和联合强化学习(RL)阶段,模型通过与真实和合成环境的交互提升能力。Kimi K2 在开源非思考模型中达到了最先进性能,尤其在智能体能力方面。值得注意的是,K2 在 Tau2-Bench 上获得 66.1 分,ACEBench(英文)76.5 分,SWE-Bench Verified 65.8 分,SWE-Bench Multilingual 47.3 分——在非思考设置下超越了大多数开源和闭源基线。它在编码、数学和推理任务中也表现出色,LiveCodeBench v6 得分 53.7,AIME 2025 得分 49.5,GPQA-Diamond 得分 75.1,OJBench 得分 27.1,均无需扩展思考。这些结果使 Kimi K2 成为迄今为止最强大的开源大语言模型之一,尤其在软件工程和智能体任务中。我们发布了基础模型和后训练模型检查点,以促进未来智能体智能的研究和应用。

We introduce Kimi K2, a Mixture-of-Experts (MoE) large language model with 32 billion activated parameters and 1 trillion total parameters. We propose the MuonClip optimizer, which improves upon Muon with a novel QK-clip technique to address training instability while enjoying the advanced token efficiency of Muon. Based on MuonClip, K2 was pre-trained on 15.5 trillion tokens with zero loss spike. During post-training, K2 undergoes a multi-stage post-training process, highlighted by a large-scale agentic data synthesis pipeline and a joint reinforcement learning (RL) stage, where the model improves its capabilities through interactions with real and synthetic environments. Kimi K2 achieves state-of-the-art performance among open-source non-thinking models, with strengths in agentic capabilities. Notably, K2 obtains 66.1 on Tau2-Bench, 76.5 on ACEBench (En), 65.8 on SWE-Bench Verified, and 47.3 on SWE-Bench Multilingual -- surpassing most open and closed-sourced baselines in non-thinking settings. It also exhibits strong capabilities in coding, mathematics, and reasoning tasks, with a score of 53.7 on LiveCodeBench v6, 49.5 on AIME 2025, 75.1 on GPQA-Diamond, and 27.1 on OJBench, all without extended thinking. These results position Kimi K2 as one of the most capable open-source large language models to date, particularly in software engineering and agentic tasks. We release our base and post-trained model checkpoints to facilitate future research and applications of agentic intelligence.

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 18)

阅读逐段中英对照全文 →