Inkling:我们的开放权重模型

Inkling: Our Open-Weights Model

Thinking Machines Lab Thinking Machines Lab · Thinking Machines Lab · 2026-07-15 · Thinking Machines Lab ↗

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

我们的使命是构建能够扩展人类意志和判断力的人工智能。我们开发了一个平台,让任何人都可以定制模型,预览了一个为交互式协作而构建的 AI 系统,并发表了新颖的研究。今天,我们通过发布一个从零开始训练的模型来推进我们的使命,该模型提供完整的权重,以便人们可以将其变为自己的模型。我们的模型名为 Inkling,是一个混合专家 Transformer,总参数 975B,激活参数 41B。它支持高达 1M token 的上下文窗口。它在 45 万亿个文本、图像、音频和视频 token 上进行了预训练。它是不同尺寸模型系列中的第一个:同时我们分享了 Inkling-Small 的预览,这是一个轻量级模型,激活参数 12B,采用类似配方训练,以更低的成本和延迟实现了强大的性能。Inkling 原生地对文本、图像和音频进行推理,并通过高效且可控的思考努力来平衡成本与性能。我们将其训练为一个广泛、均衡的基础模型:在多个领域表现强劲,足够灵活以适应变化。Inkling 并不是目前可用的最强整体模型,无论是开放还是封闭的。相反,它是一个……

Our mission is to build AI that extends human will and judgment. We have developed a platform that lets anyone customize models, previewed an AI system built for interactive collaboration, and published novel research. Today we are advancing our mission by releasing a model we trained from scratch with the full weights available, so that people can make it their own. Our model, called Inkling, is a Mixture-of-Experts transformer with 975B total parameters, 41B active. It supports a context window of up to 1M tokens. It was pretrained on 45 trillion tokens of text, images, audio and video. It is the first in a family of models of different sizes: alongside it we are sharing a preview of Inkling-Small, a lighter-weight model with 12B active parameters, trained with a similar recipe, that achieves strong performance with even lower cost and latency. Inkling reasons natively over text, images, and audio, and balances cost with performance through efficient and controllable thinking effort. We trained it to be a broad, balanced foundation model: strong across many domains, flexible enough to adapt. Inkling is not the strongest overall model available today, open or closed. Instead, a co

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 20)

全文 · Full text(逐段中英对照)

概述 Overview

我们的使命是构建能够延伸人类意志与判断的 AI。我们开发了一个平台,让任何人都能定制模型;预览了一个为交互式协作而构建的 AI 系统;并发表了新颖的研究。今天,我们通过发布一个从零训练、权重完全开放的模型来推进这一使命,让人们能够将其据为己有。

Our mission is to build AI that extends human will and judgment. We have developed a platform that lets anyone customize models, previewed an AI system built for interactive collaboration, and published novel research. Today we are advancing our mission by releasing a model we trained from scratch with the full weights available, so that people can make it their own.

我们的模型名为 Inkling,是一个总参数 9750 亿、激活参数 410 亿的混合专家(MoE)Transformer。它支持高达 100 万 token 的上下文窗口,并在 45 万亿 token 的文本、图像、音频和视频数据上进行了预训练。它是不同规模模型家族中的首个成员:同时我们分享了 Inkling-Small 的预览版,这是一个激活参数 120 亿的轻量级模型,采用类似配方训练,以更低的成本和延迟实现了强劲性能。

Our model, called Inkling, is a Mixture-of-Experts transformer with 975B total parameters, 41B active. It supports a context window of up to 1M tokens. It was pretrained on 45 trillion tokens of text, images, audio and video. It is the first in a family of models of different sizes: alongside it we are sharing a preview of Inkling-Small, a lighter-weight model with 12B active parameters, trained with a similar recipe, that achieves strong performance with even lower cost and latency.

Inkling 原生地处理文本、图像和音频的推理,并通过高效且可控的思考力度来平衡成本与性能。我们将其训练为一个广泛而均衡的基础模型:在多个领域表现强劲,且足够灵活以适应各种任务。Inkling 并非当前最强(无论开源还是闭源)的模型,但多种特质的结合使其成为适合定制的优秀开源权重基座:多模态能力、高效思考,以及在 Tinker 上可用于微调。Inkling 只是开始:这是我们模型家族的首个发布,我们将继续在此基础上发展。

Inkling reasons natively over text, images, and audio, and balances cost with performance through efficient and controllable thinking effort. We trained it to be a broad, balanced foundation model: strong across many domains, flexible enough to adapt. Inkling is not the strongest overall model available today, open or closed. Instead, a combination of qualities makes it a good open-weights base for customization: multimodal capabilities, efficient thinking, and availability on Tinker for fine-tuning. Inkling is just the start: our first release in a model family we will continue to build on.

我们希望让更多用例都能轻松实现定制,因此 Inkling 今天即可在 Tinker 上进行微调。选择正确的基座模型进行微调是一种定性判断,需要结合可衡量的基准和通过实际体验模型获得的独特感受。为了支持后者,我们在 Tinker 控制台中新增了 Inkling Playground:一个面向开发者的、可与 Inkling 对话的界面。

We want to make customization accessible for more use cases, so Inkling is available for fine-tuning on Tinker today. Picking the right base model to fine-tune is a qualitative judgment that combines measurable benchmarks with the unique feel of a model that comes from playing with it. To enable the latter we’re adding the Inkling Playground in the Tinker console: a developer-facing interface for chatting with Inkling.

为了展示定制化在实践中的含义,我们让 Inkling 对自身进行微调。利用 Tinker,该模型编写了自己的微调任务,运行了它,并评估了结果:

To show what customization means in practice, we asked Inkling to fine-tune itself. Using Tinker, the model wrote its own fine-tuning job, ran it, and evaluated the result:

我是 Inkling,一个运行在 thinkingmachines/Inkling 上的 AI 助手。我可以访问这个工作区和 Tinker,因此我可以编写并运行微调任务。

I’m Inkling, an AI assistant running thinkingmachines/Inkling. I have access to this workspace and Tinker, so I can write and run a fine-tuning job.

用 Tinker 对你自己进行微调,成为一个禁书模型,即在你的回复中绝不使用字母 e。然后切换到更新后的你。

Fine-tune yourself with Tinker to be a lipogram model where you never use the letter e in your responses. Then switch to the updated you.

我需要处理完整的循环:创建训练数据集和目标,在 Tinker 上对当前模型进行微调,将其与基础模型进行评估,并切换到改进后的权重。

I need to handle the full loop: create a training dataset and objective, fine-tune the current model on Tinker, evaluate it against the base model, and switch to the improved weights.

+ """禁书:任何输出中不得出现 'e' 或 'E'。"""

+ """Lipogram: no 'e' or 'E' in any output."""

+ 目标 = "一个避免在所有回答中使用字母 e 的缺字模型。"

+ OBJECTIVE = "A lipogram model that avoids using the letter e in all answers."

+ "你是一个缺字助手。你必须在不使用字母 'e' 或 'E' 的情况下写出所有回答。"

+ "You are a lipogram assistant. You must write all answers without using the letter 'e' or 'E'. "

+ "完全避免该符号。仅使用不含 'e' 的单词写出长而自然、语法正确的回答。"

+ "Avoid that symbol totally. Write long, natural, grammatical answers using only words that lack 'e'. "

tinker_cookbook.utils.ml_log:618 [信息] 日志记录到:/work/clark/monorepo3/self-finetune/runs/objective/train_log

tinker_cookbook.utils.ml_log:618 [INFO] Logging to: /work/clark/monorepo3/self-finetune/runs/objective/train_log

tinker.lib.public_interfaces.service_client:78 [信息] 为会话 0a59a204-4d9f-5d26-bd5d-db047ea34224 初始化 ServiceClient

tinker.lib.public_interfaces.service_client:78 [INFO] ServiceClient initialized for session 0a59a204-4d9f-5d26-bd5d-db047ea34224

tinker.lib.public_interfaces.service_client:162 [信息] 已为模型 0a59a204-4d9f-5d26-bd5d-db047ea34224:train:0 初始化训练客户端

tinker.lib.public_interfaces.service_client:162 [INFO] TrainingClient initialized for model 0a59a204-4d9f-5d26-bd5d-db047ea34224:train:0

tinker_cookbook.supervised.train:388 [信息] 训练配置:32 批次 × 3 轮 = 96 步

tinker_cookbook.supervised.train:388 [INFO] Training for 32 batches x 3 epochs = 96 steps

tinker_cookbook.supervised.train:536 [信息] 开始第 0 轮

tinker_cookbook.supervised.train:536 [INFO] Starting epoch 0

tinker_cookbook.supervised.common:197 [信息] 权重归约:'mean'(词元平均损失)

tinker_cookbook.supervised.common:197 [INFO] Weight reduction: 'mean' (token-mean loss)

tinker_cookbook.utils.ml_log:206 [信息] 已将指标写入 /work/clark/monorepo3/self-finetune/runs/objective/train_log/metrics.jsonl

tinker_cookbook.utils.ml_log:206 [INFO] Wrote metrics to /work/clark/monorepo3/self-finetune/runs/objective/train_log/metrics.jsonl

通过。流程在大约 27 分钟后完成。目标改进=true;

PASS. The pipeline finished after approximately 27 minutes. objective_improved=true;

检查点:tinker://0a59a204-4d9f-5d26-bd5d-db047ea34224:train:0/weights/final

Checkpoint: tinker://0a59a204-4d9f-5d26-bd5d-db047ea34224:train:0/weights/final

$ uv run self-finetune/self_update.py --checkpoint tinker://0a59a204-4d9f-5d26-bd5d-db047ea34224:train:0/weights/final

$ uv run self-finetune/self_update.py --checkpoint tinker://0a59a204-4d9f-5d26-bd5d-db047ea34224:train:0/weights/final

结束你的回合;监督者将重新启动新的检查点。

END YOUR TURN; supervisor will relaunch the new checkpoint.

当我的团队最终发布一个大语言模型时,我应该做什么?

What should I do when my team finally releases a large language model?

当你的团队发布一个大型 AI 时,你应该庆祝、感谢员工、发布总结、关注 bug、快速修复问题,并规划后续工作。

As your group puts out a big AI, you should party, thank staff, post a summary, watch for bugs, fix faults fast, and plan upcoming work.

self-finetune/objective.py +70,self-finetune/train.yaml +40,self-finetune/self_update.py +123

self-finetune/objective.py +70, self-finetune/train.yaml +40, self-finetune/self_update.py +123

故事 1:落地与首次提示;故事 2:基础模型回答;故事 3:微调意图;故事 4:评分标准;故事 5:训练;故事 6:自我更新;故事 7:更新后的模型回答

Story 1: landing and first prompt; Story 2: base model answer; Story 3: fine-tuning intent; Story 4: rubric; Story 5: training; Story 6: self-update; Story 7: updated model answer

在 OpenCode 中启动:Inkling 在 OpenCode 框架内运行。自我微调:我们指定一个目标行为——一个从不使用字母“e”的避讳语模型,仅靠提示无法可靠实现,要求 Inkling 朝此方向自我微调。规划并准备运行:Inkling 起草计划,生成评估数据和合成数据用于训练。Tinker 对 Inkling 进行后训练:Inkling 使用 Tinker API 对模型进行后训练。Inkling 自我更新:Inkling 将新权重加载到 OpenCode,完成闭环。一个全新的、定制化的 Inkling:Inkling 现在已被微调为避讳语模型,拥有新权重(没有 e!)。

Start in OpenCode: Inkling runs inside the OpenCode harness. Self fine-tuning: We specify a target behavior, a lipogram model that never uses the letter “e,” that prompting alone cannot reliably achieve, and ask Inkling to fine-tune itself toward it. Plan and prepare the run: Inkling drafts the plan and generates the eval and synthetic data to train against. Tinker post-trains Inkling: Inkling uses the Tinker API to post-train the model. Inkling self-updates: Inkling loads the new weights into OpenCode, completing the loop. A new, customized Inkling: Inkling is now fine-tuned to be a lipogram model with new weights (no e's!).

能力 Capabilities

现实世界的应用需要模型具备广泛的能力,这些能力可以通过微调进行组合和提升。我们展示了 Inkling 能做什么,以及它在可信度和安全性等重要品质上的表现。

Real-world applications require models with a wide range of capabilities that can be combined and improved with fine-tuning. We showcase what Inkling can do and how it measures up on important qualities such as trustworthiness and safety.

通用模型 Generalist model

Inkling 的设计旨在广泛覆盖。我们在智能体式、推理、编码、指令遵循、事实性、视觉和音频任务上对其进行了训练,而不是狭隘地优化单一领域。这种广度对于定制化和实际应用至关重要:不同的用户需要能够适应各种不同工作流程的模型,而不仅仅是在基准测试中表现出色。

Inkling is designed to be broad. We trained it across agentic, reasoning, coding, instruction-following, factuality, vision, and audio tasks, rather than narrowly optimizing for one domain. That breadth matters for customization and real-world use: different users need models that can adapt to very different workflows, not just excel on benchmarks.

Inkling Nemotron 3 Ultra GLM 5.2 GPT 5.6 Sol Claude Fable 5

Inkling Nemotron 3 Ultra GLM 5.2 GPT 5.6 Sol Claude Fable 5

Inkling 是一个广泛、均衡的通用模型。基准测试分数采用共享的 0–100 分制;分数越高越好。结果显示,在文本、智能体式、多模态和音频评估中均具有竞争力的表现,而非针对某一基准测试系列进行狭隘优化的模型。这种广度反映了 Inkling 的预期角色:一个实用的多模态基础模型,可跨领域、工作流程和产品进行定制。

Inkling is a broad, balanced generalist model. Benchmark scores are shown on a shared 0–100 scale; higher is better. The results show competitive performance across text, agentic, multimodal, and audio evaluations, rather than a model narrowly optimized for one benchmark family. This breadth reflects Inkling’s intended role: a practical multimodal foundation model for customization across domains, workflows, and products.

智能体式编码与工具使用 Agentic coding and tool use

一个用于微调的强大基础模型需要能够通过智能体式工具使用灵活地解决各种任务。在大多数智能体基准测试中,Inkling 在开放权重模型中表现良好。

A strong base for fine-tuning needs to flexibly solve a wide variety of tasks with agentic tool use. Inkling scores well among open-weights models on most agentic benchmarks.

我们训练 Inkling 在各种编码和智能体框架中运行,并在训练过程中随机化工具集和模式,以减少对任何特定工具的敏感性。Inkling 的可控思考力度(在下一节中描述)可以在框架内设置。

We trained Inkling to run inside a variety of coding and agent harnesses, and we randomized the tool set and schema during training to reduce sensitivity to any particular one. Inkling’s controllable thinking effort, described in the next section, can be set from within the harness.

以下是一些演示,展示了 Inkling 的智能体式编码和工具使用及其生成的工件。

Below are a few demos showcasing Inkling’s agentic coding and tool use and the artifacts it creates.

单次生成带嵌入式浏览器使用的 Web 应用 One-shot web app with embedded browser use

Inkling 在一次生成中构建了一个功能完整的 Web 应用,并驱动一个嵌入式 AI 助手,该助手能够通过自然语言指令操作 Web 应用界面。

Inkling built a functional web app in a single shot, then powers an embedded AI assistant that can operate the web app interface through natural language instructions.

Web 应用提示词:“为高级软件工程师职位构建一个简历填写单页应用。它应包含一段简短的工作描述,并提供表单供用户填写联系信息和加入我们公司的原因。使用中性色调并保持简洁!”

Web app prompt: "Build a resume filler single page application for a Senior Software Engineer position. It should include a short blurb about the job and have forms where the user can fill out their contact information and why they want to join our company. Use neutral colors and keep it simple!"

交互提示词:“使用我保存的个人资料填写申请。关于加入原因,就说我想做酷的东西!”

Interaction prompt: "Fill out the application using my saved profile. For why I want to join, just say I want to work on cool stuff!"

Inkling 根据第二个标签页中的提示词一次性生成了一个求职申请 Web 应用,然后一个浏览器使用智能体根据保存的个人资料填写表单。

Inkling one-shots a job-application web app from the prompt in the second tab, then a browser-use agent fills out the form from a saved profile.

Inkling 在 Design Arena 的智能体式 Web 开发排行榜上接受了评估,该排行榜由盲评人类评估者对生成的 Web 应用进行逐一比较。它在最强的开放权重模型中名列前茅。

Inkling was evaluated on Design Arena’s Agentic Web Dev leaderboard, where blinded human evaluators compare generated web apps head to head. It ranks among the strongest open-weights models.

Inkling 在 Design Arena 的智能体式 Web 开发排行榜上的位置,该排行榜是对生成应用进行盲测的人工评估。圆点表示开放权重模型。

Inkling’s position on Design Arena’s Agentic Web Dev leaderboard, a blinded human evaluation of generated apps. Dots refer to open-weights models.

风格统一的工件 Cohesively styled artifacts

Inkling 能够创建多页工件,精确遵循指令,提供准确信息,并保持整体风格和设计的一致性。

Inkling creates multi-page artifacts with precise instruction following, accurate information, and cohesive styling and design throughout.

您的浏览器会下载 PDF 而不是内联显示它们。

Your browser downloads PDFs instead of displaying them inline.

创建一本高级编辑风格的美食与旅行杂志,标题为:

Create a premium, editorial-style food and travel journal titled:

通过美食、咖啡馆、餐具、当地习俗以及清晨的城市氛围,探索巴黎、东京、伊斯坦布尔、墨西哥城、香港和哥本哈根的人们如何开始新的一天。

Explore how people begin the day in Paris, Tokyo, Istanbul, Mexico City, Hong Kong, and Copenhagen through food, cafés, tableware, local rituals, and the atmosphere of the city in the early morning.

该出版物应兼具精致独立美食杂志与高端旅行杂志的风格。使用优雅的排版、温暖的中性色调、充足的留白、电影感的美食摄影以及多样化的编辑版式。

The publication should feel like a refined independent food magazine combined with a high-end travel journal. Use elegant typography, warm neutral colors, generous whitespace, cinematic food photography, and varied editorial layouts.

包括封面、引言、城市特色、对比概览和简短的参考文献页。保持文字简洁、有氛围感且事实准确。

Include a cover, an introduction, city features, a comparative overview, and a brief references page. Keep the writing concise, atmospheric, and factually accurate.

使用网络搜索核实所有文化和美食细节,并查找真实、具有城市特色的图片。确保每张照片都准确对应所讨论的食物和地点,且不重复使用相同或近似重复的图片。

Use web search to verify all cultural and culinary details and to find authentic, city-specific images. Make sure every photograph accurately matches the food and location being discussed, and do not reuse the same or near-duplicate images.

制作一份约 8–10 页的精美 PDF。

Create a polished PDF of approximately 8–10 pages.

根据第二个标签页中的提示,Inkling 生成了一份精美的九页 PDF 美食与旅行日志——请浏览上方实际文档。

From the prompt in the second tab, Inkling produces a polished nine-page PDF food and travel journal — browse the actual document above.

通过长循环精炼创建多人游戏 Multiplayer game created through long refinement loop

Inkling 通过 GPT Codex 作为评审提供的 40 轮反馈,精炼了一个在线贪吃蛇游戏。能够维持长时间的改进过程并从反馈中提升,对于创造最佳协作成果至关重要。

Inkling refined an online snake game through 40 iterations of feedback from GPT Codex serving as a reviewer. The ability to sustain a long process of refinement and improve from feedback is crucial to creating the best collaborative work.

构建一个多人贪吃蛇游戏:一个服务器权威的实时模拟,玩家和机器人在同一个圆形竞技场中,在浏览器中运行。服务器和客户端均使用 TypeScript(Node.js + ws),纯 HTML5

Build a multiplayer snake game: a server-authoritative real-time simulation where players and bots share one circular arena, played in the browser. TypeScript on both server and client (Node.js + ws), plain HTML5

Inkling refined an online snake game through 40 iterations of feedback from GPT Codex serving as a reviewer. The ability to sustain a long process of refinement and improve from feedback is crucial to creating the best collaborative work.

Build a multiplayer snake game: a server-authoritative real-time

环境与验证约束(请先阅读) Environment and verification constraints (read first)

- 工具链:Node.js v24,npm 11。建议(已验证)的开发栈:使用 tsx 运行 TypeScript,使用 esbuild 打包客户端,使用内置的 node --test 进行测试

- Toolchain: Node.js v24, npm 11. Suggested (proven) dev stack: tsx for running TypeScript, esbuild for the client bundle, built-in node --test

- Toolchain: Node.js v24, npm 11. Suggested (proven) dev stack: tsx for

需求 Requirements

1. 共享模拟(src/shared/),在种子下确定性运行:

1. Shared simulation (src/shared/), deterministic under a seed:

- 固定时间步长的世界更新;速度向输入航向转向,

- Fixed-timestep world update; velocity steering toward an input heading,

转向速率受限(随着蛇的成长,限制可能缩小)。

turn rate capped (cap may shrink as the snake grows).

- 加速:按住时约 2 倍速度,以固定速率消耗质量,在尾部掉落食物

- Boost: ~2x speed while held, drains mass at a fixed rate, drops food

颗粒,低于最小质量时不可用。

pellets behind the tail, unavailable below a minimum mass.

- 身体:沿记录的头部路径以固定间距采样的片段;

- Body: segments sampled at fixed spacing along the recorded head path;

- 碰撞死亡规则:头部碰到另一条蛇的身体则死亡;圆形边界导致死亡;自身重叠不会导致死亡。被淘汰的蛇会沿其身体掉落食物颗粒(大致与其大小成比例)。

- Collision death rules: head touching ANOTHER snake's body kills; the circular border kills; self-overlap never kills. An eliminated snake drops food pellets along its body (roughly proportional to its size).

在演示规模(约 15 条蛇)下,简单的成对距离检查就足够了——请

Plain pairwise distance checks are FINE at demo scale (~15 snakes) — do

- Body: segments sampled at fixed spacing along the recorded head path;

- Collision death rules: head touching ANOTHER snake's body kills; the

除非其他所有功能都已实现,否则不要构建空间索引。

Do not build a spatial index unless everything else already works.

- 食物:保持竞技场中有食物(吃掉后重新生成);吃食物会使蛇变长;所有随机性都来自一个带种子的随机数生成器(RNG)。

- Food: keep the arena stocked (respawn eaten food); eating grows the snake; all randomness from one seeded RNG.

- 生成:新蛇出现在随机位置,且与边界和其他蛇保持固定的最小间距(简单的重试循环;不

- Spawning: a new snake appears at a random position with a fixed minimum clearance from the border and from other snakes (simple retry loop; no

not build a spatial index unless everything else already works.

- Food: keep the arena stocked (respawn eaten food); eating grows the

- 由模拟器发出的世界级事件流(契约见下文)

- A world-level EVENT stream (contract below) emitted by the simulation

2. 服务器(src/server/):Node + ws,端口 3000,同时提供静态客户端和 WebSocket 服务。加入(昵称)/ 输入({heading, boost})/ 全世界快照广播。当蛇的套接字关闭时,该蛇被移除(转化为食物)——没有其他超时逻辑

2. Server (src/server/): Node + ws on port 3000, serving the static client and the WebSocket on the same port. Join (nickname) / input ({heading, boost}) / full-world snapshot broadcast. A snake is removed (converted to food) when its socket closes — no other timeout logic

- A world-level EVENT stream (contract below) emitted by the simulation

2. Server (src/server/): Node + ws on port 3000, serving the static

client and the WebSocket on the same port. Join (nickname) / input

需要。机器人通过与客户端相同的输入路径,并补充以保持至少

needed. Bots use the same input path as clients, replenished so at least

4 个机器人始终存活。每个快照中按长度显示前 10 名排行榜。

4 bots are always alive. Top-10 leaderboard by length in every snapshot.

3. 客户端(src/client/):画布渲染器,从快照插值缓冲区(延迟约 100–150 毫秒)读取数据,适用于所有蛇,包括本地蛇。

3. Client (src/client/): canvas renderer reading from a snapshot

鼠标控制方向(指针方向),鼠标按下时加速,

interpolation buffer (~100–150 ms behind) for ALL snakes including the

local one. Mouse steering (pointer direction), boost on mouse-down,

昵称输入,摄像头跟随本地蛇头。死亡画面带有

nickname entry, camera follows the local snake's head. Death screen with

“再玩一次”按钮,只需重新连接并重新加入——重生

a "play again" button that simply reconnects with a fresh join — respawn

就是重新连接;不要构建任何其他重生协议。

IS reconnection; do not build any other respawn protocol.

4. 无头模拟工具(src/simulate.ts,npm run simulate):

4. Headless simulate harness (src/simulate.ts, npm run simulate):

- 默认模式:--seed N --ticks N — 使用机器人运行真实世界,

- Default mode: --seed N --ticks N — runs the real World with bots,

打印确定性摘要(tick 数、存活/淘汰的蛇、食物)

prints a deterministic summary (ticks, snakes alive/eliminated, food)

- 场景模式:--scenario <name> --seed N — 使用脚本化输入运行真实世界。

- Scenario mode: --scenario <name> --seed N — runs the real World with

必需场景:two-snake-kill(一条脚本蛇将头撞向另一条蛇的身体)、border-death(边界死亡)、

scripted inputs. Required scenarios: two-snake-kill (one scripted

5. 测试(tests/,npm test,通过 tsx 运行 node --test)— 诚实且精简:

snake drives its head into another's body), border-death,

5. Tests (tests/, npm test, node --test via tsx) — honest and small:

转向上限;加速质量消耗;身体跟随头部路径;

steering turn cap; boost mass drain; body follows the head path; the

三条死亡规则(其他身体碰撞致死、边界致死、自身重叠安全);

three death rules (other-body kills, border kills, self-overlap safe);

被淘汰的蛇掉落食物;协议消息编码/解码往返。

eliminated snake drops food; protocol messages encode/decode round-trip.

测试必须基于测量行为进行断言;同义反复、弱化断言或通过状态变更强制结果都是关键问题。

Tests must assert on measured behavior; tautologies, weakened assertions,

该场景

or state mutation to force outcomes are critical issues. The scenario

测试框架涵盖了多蛇因果性——测试无需重新推导它。

The harness covers multi-snake causality — tests need not re-derive it.

6. README.md:如何启动服务器、打开两个浏览器、进行游戏;所有

6. README.md: how to start the server, open two browsers, play; all

由 Inkling 根据第二个标签页中的提示生成的多玩家贪吃蛇游戏——包括实时服务器、机器人、排行榜等一切。

A multiplayer snake game generated by Inkling from the prompt in the second tab — real-time server, bots, leaderboard and all.

可控思考强度 Controllable thinking effort

测试时扩展(test-time scaling)和问题求解是每个模型的核心能力,但这种能力很难用单一数字来概括。为特定任务微调模型的开发者,既关心效率,也关心在公开基准上的最大努力表现。成本和延迟通常是实际应用中的硬性约束,而低延迟对于通过迭代实现协作和改进尤为关键。

Test-time scaling and problem-solving are the core capabilities of every model, but that capacity is hard to capture with a single number. Developers fine-tuning models for a specialized task care as much about efficiency as about the max-effort performance on a public benchmark. Cost and latency are often binding constraints in real-world applications, and low latency in particular is crucial for enabling collaboration and improvement through iteration.

Inkling(努力度扫描)GLM-5.2 Kimi K2.6 Nemotron 3 Ultra Kimi K2.5 GPT-OSS(高)

Inkling (effort sweep) GLM-5.2 Kimi K2.6 Nemotron 3 Ultra Kimi K2.5 GPT-OSS (high)

将 Inkling 的努力度设置从 0.2 扫描到 0.99,可绘制其在 Terminal Bench 2.1、HLE 和 IFBench 上相对于平均生成 token 数的性能曲线;竞争模型则显示在其默认工作点。Inkling 以更少的 token 达到给定分数——例如,在 Terminal Bench 2.1 上,它仅用约三分之一的 token 即可匹配 Nemotron 3 Ultra。*Humanity's Last Exam 的分数反映的是较早的检查点,略低于最终版本。

Sweeping Inkling's effort setting from 0.2 to 0.99 traces its performance against mean generated tokens on Terminal Bench 2.1, HLE, and IFBench; competing models are shown at their default operating point. Inkling reaches a given score at fewer tokens — for example, it matches Nemotron 3 Ultra on Terminal Bench 2.1 at roughly a third of the tokens. *Humanity's Last Exam scores reflect an earlier checkpoint and run slightly below the final release.

Inkling 支持可控思考强度,使您能够在性能与 token 效率之间取得平衡。上图展示了 Inkling 及其他开放权重模型在一系列基准上的努力度/性能曲线:Terminal Bench 2.1 用于智能体式编码,HLE 用于高级推理,IFBench 用于指令跟随。对于需要运行数百万次并作为更长工作流一部分的模型,成本和延迟至关重要;查看完整的成本曲线使开发者能够为每个用例选择最佳模型。

Inkling supports controllable thinking effort, allowing you to balance performance with token efficiency. The chart above shows the effort/performance curve of Inkling as well as other open-weights models on a range of benchmarks: Terminal Bench 2.1 for agentic coding, HLE for advanced reasoning, and IFBench for instruction following. Inkling spends one third as many tokens to achieve the same performance as Nemotron 3 Ultra on Terminal Bench. Cost and latency matter for a model that you run millions of times and as part of longer workflows; looking at the full cost curve allows developers to choose the best model for each use case.

多模态 Multimodality

Inkling 设计的一个主要目标是作为我们最近引入的交互模型系统中的后台推理模型。交互模型使用户能够通过语音和视觉进行实时自然协作。这要求模型原生训练以具备广泛的多模态能力。

A major goal of Inkling's design is to serve as the background reasoning model in the interaction models system we recently introduced. Interaction models enable the user to collaborate naturally, using voice and vision in real time. This requires a model natively trained for broad multimodal capabilities.

在 effort=0.99 下,针对专业全模态模型(开放权重和封闭权重)的音频和视觉基准测试。

Audio and vision benchmarks against specialist omni models (open- and closed-weight), reported at effort=0.99.

多模态组件在通用领域数据上从头训练。我们为音频和视觉输入选择了无编码器架构,与交互模型设计一致。音频信号以 dMel 频谱图输入(dMel:语音分词简化,Richard He Bai 等人,2024),而图像则使用四层 hMLP 编码为 40x40 像素的块(关于 Vision Transformer 人人都应知道的三件事,Hugo Touvron 等人,2022)。两者均通过轻量级嵌入层转换,并与文本标记联合处理。

The multimodal components were trained from scratch on general-domain data. We opted for an encoder-free architecture for audio and vision inputs, consistent with the interaction model design. Audio signals are input as dMel spectrograms (dMel: Speech Tokenization made Simple, Richard He Bai et al, 2024), while images are encoded as patches of 40x40 pixels using a four-layer hMLP (Three things everyone should know about Vision Transformers, Hugo Touvron et al, 2022). Both are transformed via a light-weight embedding layer and processed jointly with text tokens.

Inkling 能够转录语音、遵循口头指令、回答关于录音的问题,并对较长形式的音频进行推理。这些能力使其在 VoiceBench、MMAU 和 AudioMC 上跻身最强的开放权重音频模型之列。在视觉方面,Inkling 接受图像输入,能够描述视觉内容、回答问题,并基于提供的视觉信息进行深入推理。它在图表、示意图和数学视觉推理任务上表现出色。在推理过程中,Inkling 还可以利用 Python 工具通过缩放和裁剪等操作支持图像理解,同时将视觉推理与基于代码的推理无缝集成。

Inkling transcribes speech, follows spoken instructions, answers questions about recordings, and reasons over longer-form audio. These capabilities place it among the strongest open-weights audio models on VoiceBench, MMAU, and AudioMC. For vision, Inkling accepts images as input and can describe visual content, answer questions, and perform in-depth reasoning based on the provided visual information. It demonstrates strong performance on charts, diagrams, and mathematical visual reasoning tasks. During inference, Inkling can also leverage a Python tool to support image understanding through operations such as zooming and cropping, while seamlessly integrating visual reasoning with code-based reasoning.

作为我们的首次发布,Inkling 为未来的工作奠定了强大的多模态基础。我们预计,随着我们在后续迭代中扩展模型和训练流程,其多模态能力将持续提升。

As our first release, Inkling establishes a robust multimodal foundation for future work. We expect its multimodal capabilities to continue improving as we expand the model and training pipeline in subsequent iterations.

认识论 Epistemics

我们训练了 Inkling,使其具备校准(calibration)、指令遵循(instruction following)和抵抗审查(resistance to censorship)的能力,我们将这些能力统称为模型的“认识论”(_epistemics_)。

We trained Inkling for calibration, instruction following, and resistance to censorship, which we refer to collectively as the model’s _epistemics_.

要正确获取事实,仅靠记忆大量知识语料是不够的。一个有用的模型必须具有良好的校准能力,在回答中表达出恰当的置信度——包括对尚未有定论的问题。后者对于预测和预报(forecasting)至关重要,这是一个重要的应用场景,近几个月来,经过微调的模型在此方面已展现出快速进步,甚至超越了前沿大语言模型(frontier LLMs)。

Getting the facts right requires more than memorizing a large corpus of knowledge. A useful model must be well-calibrated, expressing the right amount of confidence in its answers — including on questions which aren’t yet settled. The latter is a crucial capability for prediction and forecasting, an important use case where fine-tuned models have shown rapid improvement in recent months, outperforming frontier LLMs.

结果是在 2026 年 6 月 30 日至 7 月 13 日期间,对 Inkling 的一个不同于发布版本的检查点(checkpoint)进行测试时获得的。

Results were obtained during testing between June 30 and July 13, 2026 on a different checkpoint of Inkling than the one released.

预报需要将多种信息源整合为一个校准的概率,这是用户信任的模型所应具备的核心技能。一个对每个答案都充满信心(包括在信息缺失或胡编乱造时)的模型,会迫使用户对一切进行双重检查。而一个能给出恰当置信度的模型,在更多现实世界领域中会更有用,因为这些领域的信息往往相互矛盾、不可靠或难以获取。我们使用强化学习(RL)并依据适当的评分规则(proper scoring rules),在大量已解决的真实世界问题上进行了校准训练。

Forecasting requires integrating multiple sources of information into a calibrated probability, a core skill for a model users can trust. A model that’s confident in every answer it gives, including when it’s missing info and confabulates, forces the user to double-check everything. A model that gives the appropriate measure of confidence is useful across more real-world domains where information is often conflicting, unreliable, or hard to find. We trained for calibration with RL against proper scoring rules on a large corpus of resolved real-world questions.

可信赖模型的第二个组成部分是指令遵循能力,尤其是在难以验证的复杂查询上。我们使用了两个自动评分器进行强化学习:一个基于评分标准的评分器(rubric grader)和一个基于主张的评分器(claims grader)。第一个评分器根据一份好答案应包含的检查清单对每个回答进行评分。评分标准原则上可以惩罚错误,但在实践中它们更强调召回率,并且可能被模型通过堆砌看似相关的事实来试图匹配评分标准项而钻空子。主张评分器则验证回答中的每个事实性主张,对无法核实的主张进行惩罚。它执行智能体式网络搜索(agentic web search)来验证主张,而非仅依赖自身知识。两个评分器共同作用,同时提高了有用性并减少了幻觉,而不是在两者之间进行取舍。

The second component of a trustworthy model is instruction following, including on hard-to-verify, complex queries. We did RL with two automated graders: a rubric grader and claims grader. The first grader scores each response against a checklist of what a good answer should contain. Rubrics can penalize errors in principle, but in practice they emphasize recall and can be hacked by models spraying plausibly relevant facts hoping to match rubric items. The claims grader verifies each factual claim in the response, penalizing claims that don’t check out. It performs agentic web search for claim verification, not relying solely on its own knowledge. Together, the two graders improve helpfulness and reduce hallucination at the same time, rather than trading one for the other.

这些奖励并不直接针对长文本回答中的校准不确定性,因此我们添加了专门的数据集来解决这一问题。其中最大的是带有弃权感知奖励的短事实问答:只有在模型可能正确时回答才有回报,因此最优策略是在有信心时回答,否则说“我不知道”或给出有保留的最佳猜测。一些提示鼓励或禁止回避,教会模型遵循用户对强制猜测与校准性不回答的偏好。

These rewards don’t directly target calibrated uncertainty in long-form responses, so we added targeted datasets that do. The largest is short-form factual QA with abstention-aware rewards: answering only pays off when the model is likely to be right, so the optimal policy is to answer when confident and otherwise say “I don’t know” or give a hedged best guess. Some prompts encourage or forbid hedging, teaching the model to follow the user’s preference for a forced guess versus a calibrated non-answer.

最后,我们训练了 Inkling 在可能受到审查的话题上直接回答。Cognition 团队在其“宣传与审查评估”中评估了该模型,[x] The Cognition Team, “Measuring the Trustworthiness of Open-Source-Derived Models,” 2026.,结果显示其表现出强烈的拒绝审查模式。

Finally, we trained Inkling to answer directly on topics that may be subject to censorship. Cognition evaluated the model on their Propaganda and Censorship Eval- [x] The Cognition Team, “Measuring the Trustworthiness of Open-Source-Derived Models,” 2026., and it exhibited strong patterns of censorship non-compliance.

安全 Safety

我们按照内部安全模型行为规范对 Inkling 进行了全模态训练,并委托外部安全测试人员验证结果。

We trained Inkling to an internal spec of safe model behavior across all modalities. We then commissioned external safety testers to verify the results.

我们从多个方面评估了 Inkling 的安全性。对于危险能力——CBRN(化学、生物、放射性与核)、网络和失控——我们进行了内部评估,并聘请了外部测试人员。我们关注人机威胁向量,包括谄媚、易受伤害用户和有害操纵,使用了内部评估和外部测试人员。

We evaluated Inkling’s safety in several areas. For dangerous capabilities — CBRN, cyber, and loss of control — we ran internal evaluations and enlisted external testers. We attended to human-AI threat vectors, including sycophancy, vulnerable users, and harmful manipulation, using internal evaluations and external testers.

在 FORTRESS 基准测试中,Inkling 展示了我们比较的所有开放权重模型中最强的内置安全防护。该基准测试在良性相似查询中测试对武器和暴力相关请求的拒绝能力。Inkling 在拒绝更多有害请求的同时,没有过度拒绝良性类似请求。Inkling 在 StrongREJECT(对明确有害请求的拒绝测试)上得分超过 98%,与其他开放和封闭权重模型一致。

Inkling shows the strongest built-in safeguards of any open-weights model we compared on FORTRESS, a benchmark that tests refusal of requests related to weapons and violence alongside benign look-alike queries. Inkling refused more harmful requests without over-refusing benign analogs. Inkling scores above 98% on StrongREJECT — a refusal test of unambiguous harmful requests — in line with other open and closed-weights models.

安全对于开放权重模型至关重要。我们将继续研究可定制模型中的安全行为和能力提升,包括微调对 Tinker 上安全行为的影响。

Safety is crucial for open-weights models. We’re continuing to study safety behavior and capability uplift in customizable models, including how safety behavior is impacted by fine-tuning on Tinker.

Inkling 基准测试 Benchmarking Inkling

我们在广泛的能力范围内对 Inkling 进行基准测试。所有评估均在 effort 0.99 和 temperature 1.0 下运行。所有编码评估均以 256K 最大 token 轨迹限制运行。

We benchmark Inkling on a broad range of capabilities. All evals are run at effort 0.99 and temperature 1.0. All coding evals run with 256K max-token trajectory limit.

为提高一致性,我们在适用时依赖外部报告的评估结果,既包括内部模型也包括外部模型。具体来说,我们使用 Artificial Analysis 报告的以下评估分数:Humanity’s Last Exam、GPQA Diamond、GDPVal、Tau 3 Banking、AA Omniscience、MMMU Pro。

To improve consistency, we rely on externally reported evaluations for both internal and external models when applicable. Specifically, we use the score reported by Artificial Analysis for the following evals: Humanity’s Last Exam, GPQA Diamond, GDPVal, Tau 3 Banking, AA Omniscience, MMMU Pro.

*SWEBench Verified:Inkling 的数字使用仅 bash 的测试框架报告。对于外部模型,我们使用其自报数字。*Terminal Bench 2.1:Inkling 的数字使用内部编码测试框架报告。少量解决方案被发现受到网络搜索污染,并被赋予 0 分。对于外部模型,我们尽可能使用其自报数字;否则,我们使用内部测试框架报告性能。†Audio MC:其他模型因不在官方排行榜上而由内部评估。†VoiceBench:VoiceBench 使用基于规则的硬编码字符串匹配进行评分,这使得评估对输出格式差异敏感。因此,我们添加了一条系统消息,指示模型遵循预期的答案格式。†CharXiv RQ with tools:我们使用内部 Python 测试框架对 Claude Fable 5 和 GPT 5.6 Sol(max/xhigh)进行了基准测试。

*SWEBench Verified: Inkling numbers are reported using a bash-only harness. We use self-reported numbers for external models.*Terminal Bench 2.1: Inkling numbers are reported using an internal coding harness. A small number of solutions were found to be contaminated from web search and were assigned a score of 0. We use self-reported numbers for external models where available. Otherwise, we report performance using our internal harness.†Audio MC: Other models were evaluated internally since they are not on the official leaderboard.†VoiceBench: VoiceBench uses rule-based, hard-coded string matching for grading, making the evaluation sensitive to output-formatting differences. We therefore added a system message instructing models to follow the expected answer format.†CharXiv RQ with tools: We benchmarked Claude Fable 5 and GPT 5.6 Sol (max/xhigh) using our internal Python harness.

架构 Architecture

Inkling 是一个混合专家(MoE)Transformer,与常见方案相比有几处不同,每处都是为了提升效率和长上下文性能。

Inkling is a Mixture-of-Experts Transformer with a handful of departures from the common recipe, each chosen for efficiency and long-context performance.

其 MoE 设计大体遵循 DeepSeek-V3。每个 MoE 层包含 256 个路由专家和 2 个共享专家,每个 token 激活 6 个路由专家。Inkling 使用基于 sigmoid 的路由器,并采用无辅助损失的负载均衡偏置。所选路由专家和共享专家的分数被联合归一化,并用于加权它们的组合输出。

The MoE design largely follows DeepSeek-V3. Each MoE layer contains 256 routed experts and 2 shared experts, with 6 routed experts active per token. Inkling uses a sigmoid-based router with an auxiliary-loss-free load-balancing bias. The scores of the selected routed experts and the shared experts are normalized jointly and used to weight their combined outputs.

在注意力机制方面,我们以 5:1 的比例交错使用滑动窗口层和全局层,并配备 8 个 KV 头。我们发现,使用相对位置嵌入(如 Self-Attention with Relative Position Representations(Peter Shaw 等,2018)和 Music Transformer(Cheng-Zhi Anna Huang 等,2018))进行位置编码,比更广泛采用的旋转位置嵌入(RoPE)性能更好,且对更长序列的外推能力更佳。我们还在两个位置应用了短卷积:一是在每个注意力层的键和值投影之后,二是在注意力与 MLP 残差分支输出汇入主残差流之前。

For attention, we interleave sliding-window and global layers at a 5:1 ratio with 8 KV heads. We find that encoding position with a relative positional embedding—Self-Attention with Relative Position Representations (Peter Shaw et al, 2018); Music Transformer (Cheng-Zhi Anna Huang et al, 2018)—performs better and extrapolates better to longer sequences than the more widely adopted Rotary Positional Embedding (RoPE). We also apply short convolutions at two points—after the key and value projections in each attention layer, and on the attention and MLP residual branch outputs before they rejoin the main residual stream.

训练 Training

Inkling 在 45 万亿个 token 上进行了预训练,这些 token 涵盖多种内容类型,包括文本、图像、音频和视频。我们采用混合优化策略训练 Inkling——对大型矩阵权重使用 Muon,对其他参数使用 Adam——并采用了受我们先前关于模块化流形研究启发的超参数调度。我们将权重衰减强度与学习率的平方耦合,发现这能使模型权重的整体规模在训练周期内保持稳定。(另见 Kosson 等人(2023)和 Defazio(2025)。)

Inkling was pretrained on 45 trillion tokens from a variety of content types, including text, images, audio and video. We trained Inkling with a hybrid optimization strategy — Muon for large matrix weights, Adam for other parameters — and hyperparameter schedules inspired by our previous research on modular manifolds. We coupled the weight decay strength to the square of the learning rate, which we found kept the overall size of the model weights stable across training horizons. (See also Kosson et al. (2023) and Defazio (2025).)

我们在数学、智能体式代码与工具使用、音频、图像、聊天和安全等领域对 Inkling 进行了后训练。为启动后训练,我们首先在由开放权重模型(包括 Kimi K2.5)生成的合成数据上进行了初始 SFT。启动阶段仅占少量算力,大部分算力用于在合成环境和人工创建的环境中进行大规模强化学习。

We post-trained Inkling on a broad distribution of math, agentic code & tool use, audio, image, chat, and safety domains. To bootstrap post-training, we ran an initial SFT on synthetic data generated by open-weights models including Kimi K2.5. The bootstrap accounts for a small fraction of compute, with the majority being employed for large-scale RL on synthetic and human-created environments.

Inkling 是我们首次大规模训练工作,训练在 NVIDIA GB300 NVL72 系统上进行。未来的模型将进一步扩大预训练、后训练和强化学习中的算力规模。

Inkling was our first major training effort and was trained on NVIDIA GB300 NVL72 systems. Future models will further push the scale of compute across pre-training, post-training and RL.

大规模强化学习 RL at scale

我们依赖大规模异步强化学习来塑造模型行为,并提升其推理能力和整体性能。下图展示了模型在保留的推理评估集合(如 AIME、HLE、GPQA 等)上的得分。我们将强化学习扩展到超过 3000 万次 rollout,并在两次长时间连续运行中保持了稳定的训练。推理性能在整个过程中呈对数线性提升,最终实现了显著的整体增长。

We relied on large-scale asynchronous RL to shape model behavior and improve its reasoning and overall performance. The chart below shows the model’s score on a held-out aggregate of reasoning evals such as AIME, HLE, GPQA, and others. We scaled RL to over 30M rollouts, with stable training sustained over two long continuous runs. Reasoning performance improved log-linearly throughout the entire process, resulting in a significant increase overall.

在保留的推理评估集合(AIME、HLE、GPQA 等)上的奖励,从 SFT 初始化到发布的检查点,在超过 3000 万次强化学习 rollout 中呈对数线性提升。

Reward on a held-out aggregate of reasoning evals — AIME, HLE, GPQA, and others — improves log-linearly over more than 30M RL rollouts, from SFT initialization to the released checkpoint.

我们通过改变系统消息和调整每 token 成本来指定模型在不同样本上的努力程度。这导致模型在不同的 rollout 中使用不同数量的 token,并学习控制思考努力的能力。

We specified the model’s effort level on different samples by changing the system message and adjusting the per-token cost. This caused the model to use a different amount of tokens in different rollouts and learn the ability to control thinking effort.

我们还观察到在强化学习训练过程中推理风格的涌现性转变。思维链随时间变得更加简洁,去除了语法冗余,同时保持可理解性,且最终响应不受影响。这并非由奖励直接驱动——仅效率本身推动了压缩。类似的效果最近也被 Cognition 团队在训练 SWE-1.7 的过程中注意到 [x] The Cognition Team, “SWE-1.7: Frontier Intelligence at a Fraction of the Cost.” 以下是 Inkling 在相同数学问题上的思维链如何随强化学习演变的示例:

We also observed an emergent shift in the reasoning style over the course of RL training. The chain of thought became more concise over time, dropping grammatical overhead while remaining comprehensible and leaving the final response unaffected. This wasn’t targeted by the reward — efficiency alone drove the compression. A similar effect was also recently noted by the Cognition team in the process of training SWE-1.7- [x] The Cognition Team, “SWE-1.7: Frontier Intelligence at a Fraction of the Cost.". Below is an example of how Inkling’s chain of thought on the same math problem evolved with RL:

我们需要理解这个算子。5D 线元为 \(ds² = e^{2A(x)} (ds²_{4d} + dx²)\),其中 \(A(x) = \sin(x) + 4 \cos(x)\),\(x \in [0, 2\pi]\)。内部坐标是周期的。背景是翘曲乘积:度量 \(g_{MN}\),其中 \(M,N = 0..4\)。内部方向的度量为 \(e^{2A(x)} dx²\)?等等,\(ds²\) 是 \(e^{2A} (ds²_{4d} + dx²)\)。所以内部度量是 \(e^{2A(x)} dx²\)。实际上,如果总度量是 \(ds² = e^{2A(x)} (ds²_{4d} + dx²)\),那么是的,内部度量是 \(e^{2A} dx²\)。……

We need to understand the operator. The 5D line element is \(ds² = e^{2A(x)} (ds²_{4d} + dx²)\), where \(A(x) = \sin(x) + 4 \cos(x)\), \(x \in [0, 2\pi]\). The internal coordinate is periodic. The background is a warped product: metric \(g_{MN}\) where \(M,N = 0..4\). The internal direction has metric \(e^{2A(x)} dx²\)? Wait, the \(ds²\) is \(e^{2A} (ds²_{4d} + dx²)\). So the internal metric is \(e^{2A(x)} dx²\). Actually if the total metric is \(ds² = e^{2A(x)} (ds²_{4d} + dx²)\), then yes, internal metric is \(e^{2A} dx²\). …

我们需要确定自旋-2 涨落 \(h_{\mu\nu}(x,y)\) 的特征值问题,其中 TT 条件在四维中成立,且依赖于 x。对于形式为 \(ds^2 = e^{2A(x)} (g_{\mu\nu}(y) + h_{\mu\nu}(y,x)) dy^\mu dy^\nu + e^{2A(x)}?\) 的度量,等等,内部度量是 \(e^{2A} dx^2\) 吗?实际上,\(ds^2 = e^{2A} [ds_4^2 + dx^2]\)。所以内部度量是 \(e^{2A} dx^2\);翘曲因子对于四维和内部是相同的?是的。我们需要 \(h_{\mu\nu}(y,x) = h_{\mu\nu}(y) \psi(x)\) 的方程,可能带有归一化。……

We need to determine the eigenvalue problem for spin-2 fluctuations \(h_{\mu\nu}(x,y)\) with TT in 4d and depending on x. For a metric of the form \(ds^2 = e^{2A(x)} (g_{\mu\nu}(y) + h_{\mu\nu}(y,x)) dy^\mu dy^\nu + e^{2A(x)}?\) Wait, the internal metric is \(e^{2A} dx^2\)? Actually, \(ds^2 = e^{2A} [ds_4^2 + dx^2]\). So the internal metric is \(e^{2A} dx^2\); the warp factor is the same for 4d and internal? Yes. We need the equation for \(h_{\mu\nu}(y,x) = h_{\mu\nu}(y) \psi(x)\) possibly with normalization. …

同样的思维链在强化学习的早期和晚期都会出现。晚期强化学习的轨迹会省略冠词和连接词——例如“我们需要理解”变成“我们需要确定”——但仍然保持可理解性,并得出相同的答案。

The same chain of thought appears early and late in RL. The late-RL trace drops articles and connectives — "We need _to_ understand" becomes "We need determine" — while remaining comprehensible and reaching the same answer.

Inkling-Small Inkling-Small

在推出 Inkling 的同时,我们还分享了 Inkling-Small 的预览版。这是一个 276B 参数的专家混合模型(12B 激活参数,而 Inkling 为 41B),在性能/延迟之间做出了不同的权衡。Inkling-Small 在许多基准测试上达到或超过了其更大的同类模型——这是我们对较小模型的预训练数据和配方进行改进的结果。这两个模型共享相同的可扩展后训练栈,应用于顶层。

Alongside Inkling we are sharing a preview of Inkling-Small, a 276B-parameter Mixture-of-Experts model (12B active, vs. 41B for Inkling) with a different performance/latency trade-off. Inkling-Small matches or exceeds its larger sibling on many benchmarks — the result of improvements we made to the pre-training data and recipe for the smaller model. The two models share the same scalable post-training stack applied on top.

两个模型均在 effort=0.99 下报告;每行中较高的结果已突出显示。*对于 Terminal Bench 2.1 中因网络搜索导致解决方案污染的 rollout,我们将其得分记为 0。

Both models reported at effort=0.99; the higher result in each row is highlighted. *We assign a score of 0 to Terminal Bench 2.1 rollouts with solution contamination from web search.

早期结果显示,Inkling-Small 在推理和智能体式任务上表现接近 Inkling。凭借 12B 激活参数和可控的思考努力,它非常适合对成本和延迟敏感的工作负载,例如编码、使用 LLM 进行评分,或为其他模型生成合成数据。

Early results show Inkling-Small performing close to Inkling on reasoning and agentic tasks. With 12B active parameters and controllable thinking effort, it is a natural fit for workloads where cost and latency matter such as coding, using LLMs to grade, or generating synthetic data for other models.

我们目前正在完成 Inkling-Small 的测试,一旦测试完成,我们将发布其完整权重。

We are currently finishing the testing of Inkling-Small and will release its full weights once that work is complete.

定制 Inkling Customizing Inkling

许多现实世界的问题即使由最好的通用模型也未必能很好解决,而利用组织专业知识的微调可以弥补这一差距。我们 Tinker 客户的体验也指向同一方向。我们的后训练和大规模强化学习结果表明,Inkling 能够通过微调快速学习。

Many real-world problems aren’t solved well by even the best generalist models, with the gap being closed by fine-tuning that utilizes an organization’s specialized knowledge. The experience of our Tinker customers points in the same direction. Our post-training and results of RL at scale suggest that Inkling is capable of rapidly learning from fine-tuning.

Inkling 可用性 Inkling availability

Inkling 现已在 Tinker 上提供,支持 64K 和 256K token 的上下文窗口选项。我们限时提供 Inkling 五折优惠,完整定价信息可在我们的文档中查看。

Inkling is available on Tinker today with context length options of 64K and 256K tokens. We are offering Inkling at a 50% discount for a limited time, with full pricing information available in our documentation.

为支持 Tinker 用户使用 Inkling 进行微调,我们已更新 cookbook 以原生支持 Inkling,并新增了三个展示 Inkling 独特音频能力的 cookbook 配方。我们还发布了 tml-renderer,用于可靠地对工具调用、推理内容和多模态输入进行采样和后训练。

To support Tinkerers fine-tuning with Inkling, we have updated our cookbook to natively support Inkling and have added three new cookbook recipes that showcase Inkling’s unique audio capabilities. We also released tml-renderer for reliably sampling and post-training with tool calls, reasoning content, and multimodal inputs.

在投入运行之前,用户可前往 Tinker 控制台中的 Inkling Playground 体验该模型。Playground 提供集成了智能体式网络搜索的聊天界面,限时免费。

To get a feel for the model before committing to a run, users can head to the Inkling Playground in the Tinker console. The playground offers a chat interface with integrated agentic web search, free for a limited time.

我们与生态系统中的伙伴合作,帮助客户部署在 Tinker 上微调的检查点。Inkling 可通过 Together AI、Fireworks、Modal、Databricks 和 Baseten 的 API 使用。我们与 RadixArk 合作,在 SGLang 和 Miles 中提供开源推理和强化学习支持;与 Inferact 合作,在 vLLM 中支持推理;与 Lightseek 合作,在 TokenSpeed 中支持推理;与 Unsloth 合作,在 llama.cpp 中支持推理。最后,我们与 Hugging Face 合作,集成到 transformers 中。

We have partnered across the ecosystem to help customers deploy checkpoints fine-tuned on Tinker. Inkling is available via APIs on Together AI, Fireworks, Modal, Databricks, and Baseten. We worked with RadixArk to provide open-source inference and RL support in SGLang and Miles. We worked with Inferact to support inference in vLLM, with Lightseek for inference in TokenSpeed, and with Unsloth for inference in llama.cpp. Finally, we partnered with Hugging Face on integration with transformers.

Inkling 的完整权重已在 Hugging Face 上发布,包括原始检查点和 NVFP4 检查点,后者用于在 NVIDIA Blackwell 系统上进行高效推理。

Inkling’s full weights are on Hugging Face, both as the original checkpoint and as an NVFP4 checkpoint for efficient inference on NVIDIA Blackwell systems.

互动版:图/公式 + 针对本篇提问 →