Inkling:一款带有若干惊喜的新型开放权重 975B MoE 模型

Inkling: A New Open-Weight 975B MoE with a Few Surprises

塞巴斯蒂安·拉施卡 Sebastian Raschka · · 2026-07-16 · Ahead of AI ↗

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

Thinking Machines Lab 发布了 Inkling,一款 975B 参数的开放权重混合专家(MoE)模型。每个 token 激活 41B 参数,支持高达 1,048,576 token 的上下文窗口。这些数字使 Inkling 与 Kimi K2.5 和 GLM-5.2 处于同一规模级别。然而,其架构中有几个我较少见到的细节,包括每个解码器块内的短卷积,以及用可学习的相对位置偏置替代 RoPE。

2. Local and global grouped-query attention Thinking Machines Lab released Inkling, a 975B-parameter open-weight Mixture-of-Experts (MoE) model. It activates 41B parameters per token and supports a context window of up to 1,048,576 tokens. Those numbers put Inkling in the same general size class as Kimi K2.5 and GLM-5.2. Yet the architecture has several details I have seen less often, including short convolutions inside every decoder block and a learned relative-position bias in place of RoPE.

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 6)

全文 · Full text(逐段中英对照)

概述 Overview

2. 局部与全局分组查询注意力

2. Local and global grouped-query attention

Thinking Machines Lab 发布了 Inkling,一个 9750 亿参数的开源权重混合专家(MoE)模型。它每个 token 激活 410 亿参数,并支持高达 1,048,576 个 token 的上下文窗口。

Thinking Machines Lab released Inkling, a 975B-parameter open-weight Mixture-of-Experts (MoE) model. It activates 41B parameters per token and supports a context window of up to 1,048,576 tokens.

这些数字使 Inkling 与 Kimi K2.5 和 GLM-5.2 处于相同的一般规模级别。然而,该架构有几个我较少见到的细节,包括每个解码器块内的短卷积,以及用学习到的相对位置偏置代替 RoPE。

Those numbers put Inkling in the same general size class as Kimi K2.5 and GLM-5.2. Yet the architecture has several details I have seen less often, including short convolutions inside every decoder block and a learned relative-position bias in place of RoPE.

图 1:Inkling 架构及发布时的基准比较。左图总结了 975B MoE 及其局部-全局注意力模式。基准面板使用 Thinking Machines Lab 于 2026 年 7 月 15 日发布的结果。所有 Inkling 结果均使用 0.99 的 effort 设置。更大的架构图可在 LLM 架构画廊中找到。

Figure 1: Inkling architecture and release-time benchmark comparisons. The left panel summarizes the 975B MoE and its local-global attention pattern. The benchmark panels use results published by Thinking Machines Lab on July 15, 2026. All Inkling results use an effort setting of 0.99. A larger architecture figure is available in the LLM Architecture Gallery.

41B 活跃参数规模 The 41B active footprint

Inkling 拥有 66 个解码器层,隐藏大小为 6,144。前两层使用密集前馈块,其余层包含 256 个路由专家和 2 个共享专家。每个 token 选择 6 个路由专家,且两个共享专家始终处于激活状态。

Inkling has 66 decoder layers with a hidden size of 6,144. The first two layers use dense feed-forward blocks. The remaining layers contain 256 routed experts and 2 shared experts. Each token selects 6 routed experts, and both shared experts are always active.

这相当于约 4.2% 的激活比例。作为对比,GLM-5.2 总参数为 744B,激活参数为 40B;而 Kimi K2.5 总参数为 1T,激活参数为 32B。因此,尽管 Inkling 的总参数比 GLM-5.2 多 231B,但其激活规模与 GLM-5.2 相近。

This works out to an active ratio of about 4.2%. For comparison, GLM-5.2 has 744B total and 40B active parameters, while Kimi K2.5 has 1T total and 32B active parameters. Inkling is therefore close to GLM-5.2 in active size despite having 231B more total parameters.

Inkling 也是一个原生多模态模型。图像和视频帧通过四层 hMLP 处理,而音频则使用 Thinking Machines Lab 提出的 dMel 表示。得到的表示与文本一起进入同一个解码器,模型最终输出文本。

Inkling is also a native multimodal model. Images and video frames pass through a four-layer hMLP, while audio uses the dMel representation introduced by Thinking Machines Lab. The resulting representations enter the same decoder as text, and the model produces text output.

局部与全局分组查询注意力 Local and global grouped-query attention

在 66 个解码器层中,55 层使用滑动窗口注意力,11 层使用全局注意力。这使 Inkling 呈现出重复的 5 比 1 的局部-全局模式。局部层具有 512 个 token 的窗口,使用 64 个查询头与 16 个键值头,即 4 比 1 的分组查询注意力(GQA)比率。全局层保留 64 个查询头,但仅使用 8 个键值头,将 GQA 比率提高到 8 比 1。

Of the 66 decoder layers, 55 use sliding-window attention and 11 use global attention. This gives Inkling a repeating 5-to-1 local-global pattern. The local layers have a 512-token window and use 64 query heads with 16 key-value heads, a 4-to-1 grouped-query attention (GQA) ratio. The global layers retain 64 query heads but use only 8 key-value heads, increasing the GQA ratio to 8 to 1.

不寻常之处在于位置信息。Inkling 跳过了 RoPE,而是从查询和键状态计算一个可学习的、输入依赖的相对位置偏置。它还在注意力之前对查询头和键头应用 RMS 归一化。

The unusual part is the positional information. Inkling skips RoPE and instead computes a learned, input-dependent relative-position bias from the query and key states. It also applies RMS normalization to the query and key heads before attention.

该偏置在局部层覆盖完整的 512 个 token 跨度。在全局层中,其配置范围为 1,024 个 token。更早的 token 仍然可以被注意到,但它们不会获得显式的可学习相对位置项。这个细节在仅阅读发布帖子时很容易被忽略,但在配置和 Transformers 实现中可见。

The bias covers the full 512-token span in local layers. In global layers, its configured extent is 1,024 tokens. Tokens farther back can still be attended to, but they receive no explicit learned relative-position term. This detail is easy to miss when reading only the release post. It is visible in the configuration and the Transformers implementation.

每层四个短卷积 Four short convolutions per layer

每个解码器层包含四个因果卷积,卷积核大小为 4。其中两个直接作用于键和值投影之后,另外两个处理注意力机制和 MoE 分支的输出,然后再将这些输出加入残差流。

Each decoder layer contains four causal convolutions with a kernel size of 4. Two operate directly after the key and value projections. The other two process the attention and MoE branch outputs before those outputs join the residual stream.

这些小卷积为模块提供了一条显式的路径,用于混合相邻词元的信息。该发布版本未包含隔离其效果的消融实验,因此尚不清楚 Inkling 的质量或训练稳定性在多大程度上源于这一选择。

These small convolutions give the block an explicit path for mixing information over nearby tokens. The release does not include an ablation that isolates their effect, so it is unclear how much of Inkling’s quality or training stability comes from this choice.

还有一个归一化细节。在词元嵌入查找之后、第一个解码器块之前,直接应用了一个独立的 RMSNorm。在发布的配置中启用了这种嵌入归一化,并在实现中作为一个独立操作出现。

There is one more normalization detail. A separate RMSNorm is applied directly after the token embedding lookup, before the first decoder block. This embedding normalization is enabled in the released configuration and appears as a distinct operation in the implementation.

训练与努力控制 Training and effort control

Thinking Machines Lab 报告了涵盖文本、图像、音频和视频的 45T 预训练 token。训练设置对模型的大矩阵参数使用 Muon,对其余参数使用 Adam。权重衰减与学习率的平方耦合。

Thinking Machines Lab reports 45T pre-training tokens spanning text, images, audio, and video. The training setup uses Muon for the model's large matrix parameters and Adam for the remaining parameters. Weight decay is coupled to the square of the learning rate.

后训练始于对来自开放权重教师模型(包括 Kimi K2.5)的合成数据进行监督微调。后训练的大部分算力随后投入到跨合成环境和人类创建环境的强化学习中。根据发布信息,团队收集了超过 3000 万次强化学习 rollout。

Post-training began with supervised fine-tuning on synthetic data from open-weight teacher models, including Kimi K2.5. Most of the post-training compute then went into reinforcement learning across synthetic and human-created environments. According to the release, the team collected more than 30 million RL rollouts.

一个实际成果是努力控制。系统消息告诉 Inkling 应投入多少努力,而训练期间的每 token 成本则鼓励更短或更长的回答。这对于在输出长度和基准性能之间进行权衡非常有用,而无需维护单独的检查点。

One practical outcome is an effort control. A system message tells Inkling how much effort to spend, while a per-token cost during training encourages shorter or longer answers. This is useful for trading output length against benchmark performance without maintaining separate checkpoints.

混合基准快照 A mixed benchmark snapshot

所报告的 GLM-5.2 对比展示了 Inkling 的整体基准表现。Inkling 在 IFBench 上得分为 79.8,而 GLM-5.2 为 73.3;在 SimpleQA Verified 上为 43.9 对 38.1。GLM-5.2 在无工具 HLE(40.1 对 29.7)、SWE-Bench Pro Public(62.1 对 54.3)以及 Terminal-Bench 2.1(82.7 对 63.8)上领先。

The reported GLM-5.2 comparison illustrates Inkling’s overall benchmark profile. Inkling scores 79.8 on IFBench versus 73.3 for GLM-5.2, and 43.9 versus 38.1 on SimpleQA Verified. GLM-5.2 leads on HLE without tools at 40.1 versus 29.7, SWE-Bench Pro Public at 62.1 versus 54.3, and Terminal-Bench 2.1 at 82.7 versus 63.8.

我会将这些视为发布时的快照。所有 Inkling 结果均使用 effort 设置为 0.99、温度设置为 1.0。编码评估允许轨迹长度达 256K 个 token。部分对比值来自外部报告,而 Inkling 的 Terminal-Bench 结果使用了内部测试框架。此类设置下的小差距难以解读。

I would treat these as a release-time snapshot. All Inkling results use an effort setting of 0.99 and a temperature of 1.0. The coding evaluations allow trajectories up to 256K tokens. Several comparison values come from external reports, while Inkling’s Terminal-Bench result uses an internal harness. Small gaps across such setups are hard to interpret.

发布内容还展示了 effort 扫描,而非单一操作点。在 Terminal-Bench 上,Thinking Machines Lab 报告称,Inkling 在生成约三分之一的 token 数量时达到了 Nemotron 3 Ultra 的分数。这是一个有用的结果,尽管提供商层面的吞吐量比较仍需要相同的量化、批大小、专家并行设置、注意力内核和硬件。

The release also shows effort sweeps rather than a single operating point. On Terminal-Bench, Thinking Machines Lab reports that Inkling reaches the Nemotron 3 Ultra score while generating about one-third as many tokens. This is a useful result, although a provider-level throughput comparison would still require the same quantization, batch size, expert-parallel setup, attention kernels, and hardware.

对我来说,开放性问题现在相当具体。我希望看到针对短卷积、嵌入归一化和相对位置偏置的受控消融实验。我还希望看到与 Kimi K2.5 和 GLM-5.2 直接可比的吞吐量测量。这些结果将有助于阐明 Inkling 的架构选择主要是在质量、长上下文行为、训练稳定性还是服务效率方面带来改进。

For me, the open questions are now quite concrete. I would like to see controlled ablations for the short convolutions, embedding normalization, and relative-position bias. I would also like directly comparable throughput measurements against Kimi K2.5 and GLM-5.2. Those results would clarify whether Inkling’s architecture choices mainly improve quality, long-context behavior, training stability, or serving efficiency.

互动版:图/公式 + 针对本篇提问 →