Kimi K3 Architecture Notes
打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→Kimi K3 架构图是针对昨天发布的大型开放权重模型的,附带一些观察和思考。1. 是的,它看起来相对复杂,但本质上它是他们去年发布的 Kimi Linear 模型的规模化生产版本(从 48B 扩展到 2.8T;K3 是目前最大的开放权重模型)。2. 与 Kimi Linear 相比,唯一的新组件是 LatentMoE。我在下图中省略了它,因为图已经很拥挤了,但它本质上与 Nemotron 3 Ultra 中的 LatentMoE 相同(如果你好奇,可以在我的 LLM 架构画廊中找到)。这里的想法是压缩(下投影)大型线性层,类似于多头潜在注意力。
The Kimi K3 architecture figure for yesterday’s big open-weight model release, along with some observations and thoughts. 1. Yes, it looks relatively complicated, but it’s essentially a scaled-up production version of their Kimi Linear model they released last year (scaled up from 48B to 2.8T; K3 is by far the biggest open-weight model right now) 2. The one new component compared to Kimi Linear is the LatentMoE. I omitted it in the figure below since it's already very crowded, but that's essentially the same LatentMoE as in Nemotron 3 Ultra (you can find it in my LLM Architecture Gallery if you are curious). The idea here is to compress (down-project) large linear layers similar to multi-head latent attention.
Kimi K3 架构图,针对昨天发布的大型开放权重模型,以及一些观察和思考。
The Kimi K3 architecture figure for yesterday’s big open-weight model release, along with some observations and thoughts.
1. 是的,它看起来相对复杂,但本质上它是他们去年发布的 Kimi Linear 模型的规模化生产版本(从 48B 扩展到 2.8T;K3 是目前最大的开放权重模型)
1. Yes, it looks relatively complicated, but it’s essentially a scaled-up production version of their Kimi Linear model they released last year (scaled up from 48B to 2.8T; K3 is by far the biggest open-weight model right now)
2. 与 Kimi Linear 相比,唯一的新组件是 LatentMoE。我在下图中省略了它,因为图已经很拥挤了,但它本质上与 Nemotron 3 Ultra 中的 LatentMoE 相同(如果你好奇,可以在我的 LLM 架构画廊中找到它)。其思想类似于多头潜在注意力,对大型线性层进行压缩(下投影)。
2. The one new component compared to Kimi Linear is the LatentMoE. I omitted it in the figure below since it's already very crowded, but that's essentially the same LatentMoE as in Nemotron 3 Ultra (you can find it in my LLM Architecture Gallery if you are curious). The idea here is to compress (down-project) large linear layers similar to multi-head latent attention.
3. Kimi K3 的总体趋势(类似于 Nemotron 3、DeepSeek V4 等)也是朝着更好的推理效率发展。也就是说,许多组件被替换为效率调整后的版本。例如,MoE → LatentMoE,常规注意力 → 多头潜在注意力和 Kimi Delta 注意力。(如果你好奇更多细节,我的画廊中也有简短的教程和文章)。
3. Kimi K3's overall trend (similar to Nemotron 3, DeepSeek V4, and others) is also towards better inference efficiency. That is, there are many components that replace existing components with efficiency-tweaked versions. I.e., MoE -> LatentMoE, regular attention -> multi-head latent attention and Kimi Delta Attention. (I also have short tutorials and write-ups in my gallery if you are curious about additional details).
4. 一个不属于效率调整的组件变化是注意力残差。就像 DeepSeek V4 通过 mHC(流形约束超连接)改进了残差路径,注意力残差也是一种改进残差路径的方法,但其工作方式略有不同。即 mHC 使残差路径更宽,而注意力残差(也已是 Kimi Linear 的一部分)跨层连接残差;连接本身使用注意力分数作为重要性/贡献权重。根据报告,它持续改善了验证损失和下游性能(略微),并增加了约 4% 的训练成本和 2% 的推理成本。
4. The one component change that is not an efficiency tweak is attention residuals. Like DeepSeek V4 improved the residual path with mHC (manifold-constrained Hyper-Connections), attention residuals are a way to improve the residual path, but it works a bit differently. I.e., mHC made the residual path wider. Attention residuals (also already part of Kimi Linear) connect the residuals across layers; the connection itself uses an attention score for an important/contribution weight. According to the report, it improves the validation loss and downstream performance (a bit) consistently and adds about 4% in training cost and 2% in inference cost.
5. 有趣的是,Kimi K3 去掉了所有 RoPE 层,转而全面使用 NoPE(无位置嵌入)。(同样,这继承自 Kimi Linear)。在其他架构中,最近的趋势是在局部注意力层(如滑动窗口注意力)中使用 RoPE,在全局层中使用 NoPE。有一些架构只使用 NoPE,但据我所知,这是第一个前沿级别的。
5. Interestingly, Kimi K3 got rid of all RoPE layers and uses NoPE (No Positional Embeddings) everywhere instead. (Again, this is inherited from Kimi Linear). In other architectures, the recent trend was towards RoPE in local attention layers (like sliding window attention) and NoPE in the global layers. There were a few architectures that only used NoPE everywhere, but this is the first frontier-level one as far as I know.
6. Kimi K3 现在还原生支持多模态,这很棒!
6. Kimi K3 now also has native multimodal support, which is great!
技术报告中还有其他几个有趣的训练细节,但到目前为止,架构方面就这些了。总体而言,这是一次非常出色的发布。
There are several other interesting training tidbits in the technical report, but that's it from the architecture front so far. A really great release overall.
图 1. Kimi K3 的架构及发布时的基准比较。更多细节请参见架构画廊中的 K3。
Figure 1. Kimi K3 architecture and release-time benchmark comparisons. See K3 in the architecture gallery for more details.