Kimi K3:开放前沿智能

Kimi K3: Open Frontier Intelligence

杨植麟 Zhilin Yang · · 2026-07-27 · arXiv:2607.24653 ↗ · 被引 2

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

我们推出了 Kimi K3,一个拥有 2.8 万亿参数的混合专家模型,具备 1040 亿激活参数、原生视觉能力和 100 万 token 的上下文窗口。Kimi K3 基于 Kimi Delta Attention 和 Attention Residuals 构建,这两项技术改善了序列长度和模型深度上的信息流动。结合每 token 有效激活 896 个路由专家中的 16 个的 Stable LatentMoE,以及经过优化的训练和数据方案,这些进展使得整体扩展效率相比 Kimi K2 提升了约 2.5 倍。后训练阶段的核心包括在通用、智能体和编码领域以及多种推理努力水平上的强化学习,从而实现了组合泛化和稳健的长程执行能力。在 2.8T 规模下,Kimi K3 得到了多个领域基础设施进展的支持:针对 KDA 的算法-系统协同设计、完全均衡的专家并行训练与高效内存管理、支持持续 rollout 和沙盒状态的百万 token 智能体强化学习,以及部署创新。大量评估表明,Kimi K3 在长程编码、智能体、知识、推理和视觉任务上均达到前沿水平。尽管其整体性能仍落后于最强大的专有模型,即 Claude Fable 5 和 GPT-5.6 Sol,但 Kimi K3 在评估套件中的表现持续优于其他开放和专有模型。我们发布完整的 Kimi K3 模型权重,以促进未来研究,并加速前沿智能的更广泛部署与应用。

We introduce Kimi K3, a 2.8T parameter Mixture-of-Experts model with 104 billion activated parameters, native vision capabilities, and a 1-million-token context window. Kimi K3 is built on Kimi Delta Attention and Attention Residuals, which improve information flow across sequence length and model depth. Together with Stable LatentMoE, which effectively activates 16 of 896 routed experts per token, and refined training and data recipes, these advances yield an approximately 2.5x improvement in overall scaling efficiency over Kimi K2. Post-training highlights reinforcement learning across general, agentic, and coding domains and multiple reasoning-effort levels, enabling compositional generalization and robust long-horizon execution. At 2.8T scale, Kimi K3 is supported by infrastructure advances in multiple areas: algorithm-system co-design for KDA, perfectly balanced expert-parallel training with efficient memory management, million-token agentic RL with persistent rollout and sandbox states, and deployment innovations. Extensive evaluations show that Kimi K3 achieves frontier-level performance across long-horizon coding, agentic, knowledge, reasoning, and vision tasks. While its overall performance still trails the most powerful proprietary models, namely Claude Fable 5 and GPT-5.6 Sol, Kimi K3 consistently outperforms other open and proprietary models evaluated in our suite. We release the full Kimi K3 model weights to facilitate future research and accelerate the broader deployment and adoption of frontier intelligence.

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 24)

阅读逐段中英对照全文 →