教每个人钓取令牌

Teaching Everyone to Fish for Tokens

内森·兰伯特 Nathan Lambert · Interconnects · 2026-08-17 · Interconnects ↗

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

本文探讨了开源 AI 的未来,重点关注开放权重模型与包含训练配方的完全开源模型之间的区别。文章认为,虽然开放权重模型是暂时的,但以 Olmo 等项目为代表的开源配方使公司能够构建定制模型。核心论点是,英伟达对开放模型的投资旨在创建一个自我维持的生态系统,以推动对其芯片的需求,但这一策略因训练的资金密集型特性而面临经济挑战。作者提出了两种可能的未来:一种是开放模型保持竞争力并在财务上可行,另一种是它们分化为长尾的专门应用,后者更有可能。文章总结道,开源生态系统的可持续性取决于财务反馈循环,而后训练正成为主要焦点,可能重塑传统的预训练/后训练术语。

This article examines the future of open-source AI, focusing on the distinction between open-weight models and fully open-source models with training recipes. It argues that while open-weight models are transient, the open-source recipe, exemplified by projects like Olmo, enables companies to build custom models. The core thesis is that Nvidia's investment in open models aims to create a self-sustaining ecosystem that drives demand for its chips, but this strategy faces economic challenges due to the capital-intensive nature of training. The author posits two possible futures: one where open models remain competitive and financially viable, and another where they fork into a long-tail of specialized applications, with the latter being more likely. The article concludes that the open-source ecosystem's sustainability depends on financial feedback loops, and that post-training is becoming the primary focus, potentially reshaping the traditional pretraining/post-training lexicon.

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 1)

全文 · Full text(逐段中英对照)

Nvidia 希望你构建自己的模型,而不是从 Anthropic/OpenAI 购买 Nvidia wants you building your own model, not buying from Anthropic/OpenAI.

杂项:由于我在旅行,本帖没有配音。

Housekeeping: No voiceover for this post as I'm traveling.

人们最常做的比较是,开放模型的发展与 Linux 操作系统等基础开源软件项目有何相似之处。这些类比相当清晰,但它们为开源模型生态系统的自持续性描绘了一条狭窄的前进道路,而一旦 Linux 足够强大,它就会自我实现,成为许多任务的最佳工具。开源语言模型——即只有带有完整训练配方、数据、代码等的模型——更接近开源操作系统。你使用的开放权重模型——那些只有模型权重和推理代码来运行它们的模型——更接近于你在基于它们构建的项目中安装的特定软件版本。

The oldest comparison people try to make is how what's happening with open models compares to foundational open-source software projects like the Linux operating system. There are fairly clean analogies, but they paint a narrow path forwards for the self-sustaining nature of the open-source model ecosystem, where once Linux got big enough it was going to be self-fulfilling as the best possible tool for many jobs. The open-source language model – i.e. only models that come with a full training recipe, data, code, etc. – is a closer analogue to the open-source operating system. The open weight models you use – those with just model weights and inference code to run them – are closer to specific versions of software that you install in a project built upon them.

模型权重平均而言非常短暂,但它们仍然有很长的保质期,就像许多大量使用的软件一样。这就是为什么尽管智能体行为在几年后才兴起,许多公司仍在使用基于 Llama 3 构建的工作流。开源配方,现代以我在 Ai2 帮助构建的 Olmo 模型为代表,其前身包括 EleutherAI 的 Pythia,是一个资源密集型的过程,任何公司都可以拿起、修改并按下“运行”以生成一组新的模型权重。在最好的情况下,社区还可以将数据或训练代码的改进贡献回下一个模型!这就是为什么 Nvidia 在近乎开源模型上投入如此之多——对于他们的 Nemotron 模型,他们发布了所有合法可发布的数据和训练代码等。Nvidia 希望一个无数人可以构建 token 机器的世界,这样智能就不会被垄断。这是一个对推理有巨大需求的世界,许多公司都希望购买 Nvidia 的产品。

Model weights are very transient on average, but they still have a long shelf life, as with a lot of heavily used software. It's why many companies are still using workflows built on Llama 3, despite agentic behaviors taking off years later. The open-source recipe, typified in modern times by the Olmo models I helped build at Ai2, with its predecessors like Pythia from EleutherAI, is a resource intensive process that any company can pick up, modify, and press "run" on to produce a new set of model weights. In the best cases, the community can contribute improvements in data or training code back into the next model too! This is why Nvidia is investing so much in nearly open-source models – for their Nemotron models they release all the data they legally can and the training code, etc. Nvidia wants a world where countless people can build token machines, so intelligence is not monopolized. This is a world with massive demand for inference across many companies, all of which want to buy Nvidia's offerings.

开源 AI 的未来充满变数,因为构建最好的模型极其资本密集。在行业中,构建竞争性模型的能力比许多人预期的更长时间保持可及性。许多人的默认预期是,训练模型过于昂贵,开源配方远远落后,因此围绕训练 LLM 的某个部分建立新实验室将是不可行的。

Open-source AI has a tricky future, as building the best models is extremely capital intensive. The ability to build competitive models has stayed more accessible in industry longer than many would've expected. The default expectation for many is that training models is too expensive and the open-source recipe is too far behind, so building a new lab centered on some part of training LLMs will not be tractable.

从这里出发有两种未来。第一种是“如果它有效”——如果开源配方对 Nvidia 有效,他们将创造远超构建模型成本的芯片需求(和利润)。目前据报道,Nvidia 在这项努力上花费了 260 亿美元。目前尚不清楚这是否会成功,或者 AI 的资本密集性是否会驱使越来越多的公司退出训练游戏。我们还没有看到太多迹象表明这一点。事实上,退出的公司——如 Databricks 和 01.ai——似乎是异常情况。

There are two futures from here. First is if "it works" – if the open-source recipe works for Nvidia, they'll be creating far more demand for their chips (and profits) than it costs to build the models. Right now it's reported that Nvidia is spending $26 billion on this endeavor. It's not clear if this will work, or if AI's capital intensiveness will drive more and more companies out of the training game. We haven't seen many signs of this starting. In fact, the companies bowing out – like Databricks and 01.ai – seem like anomalies.

开源生态系统在未来几年将越来越依赖 Nvidia 的资金支持。这是一个生存窗口期,几年内这种方式的利润需要回报给 Nvidia,或者另一家开源模型公司需要在其开放性上建立类似平台的财务反馈循环。这种经济回报需要与 Anthropic 和 OpenAI 的 API 所产生的利润成比例,以便在数十年的语言模型开发中保持同步。这可以通过性能上的竞争力或 AI 热潮的巨大规模来推动,使得开源模型的训练、推理和微调公司都有大量的需求。

The open-source ecosystem will become increasingly dependent on Nvidia's financing in the coming years. This is an existential window, where within a few years the profits of this approach need to return to them, or another open model company needs to cultivate platform-like financial feedback loops on their openness. This economic reward needs to be proportional to the profits generated by Anthropic and OpenAI's APIs to keep pace over decades of language model development. This can be driven by competitiveness on performance or by the AI boom just being so big that the open model training, inference, and fine-tuning companies all have vast quantities of demand.

第二种未来是,如果这两条财务上积极的道路都没有实现,开源模型将走向与领先的闭源模型不同的发展路径——一条更注重效率、可修改性、专业化等的路径。我在心里认为这是最可能的结果——开源模型仍然非常有用,但相对于闭源模型,它们填补的是一个长尾生态系统,而闭源模型在最有价值的领域(如知识工作协作、药物发现、软件工程等)拥有垄断所有权。长尾类似于企业特定的智能体,它们在本地运行,使用私有数据执行重复性业务任务。

The second future is if one of these two financially positive paths doesn't play out, open models will fork to a different development path than the leading closed models – one more focused on efficiency, modifiability, specialization, etc. I put this mentally as my most likely outcome – open models are still incredibly useful, but fill a long-tail ecosystem relative to the closed counterparts that have monopoly ownership stakes in the most valuable areas like knowledge work collaboration, drug discovery, SWE, etc. The long-tail is something like enterprise-specific agents that run on-prem with private data on repetitive business tasks.

我认为开源训练难以流行起来的部分原因是训练变得越来越复杂和抽象。当前的开源模型生态系统得益于对后训练开源模型的兴趣激增。这些人拿 DeepSeek V4 Flash、Inkling Small 或 GLM 5.X 等模型,针对他们的特定智能体任务进行微调(例如,在 Tinker 中,这是目前最流行的微调 API)。

Part of why I think this open-source training will have a hard time catching on is because training is getting more complex and more abstracted. The current open model ecosystem is buoyed by an explosion in interest in post-training open models. These people take models like DeepSeek V4 Flash, Inkling Small, or GLM 5.X and finetune them for their specific agentic tasks (e.g. in Tinker, the most popular finetuning API today).

Interconnects AI 是一个由读者支持的出版物。考虑成为订阅者。

Interconnects AI is a reader-supported publication. Consider becoming a subscriber.

在过去几年中,后训练主要指修改基础模型使其变得智能和可用的整个过程。现在正在发生一种转变,即将基础模型训练成通用智能体推理器的能力变得不透明,就像几年前的大规模预训练实践一样。这甚至可能改变已经标准了几年的预训练、中期训练、后训练术语体系。它可能会变得更接近预训练、推理训练和后训练。

Over the last few years, post-training largely referred to the whole process of modifying the base model to make it intelligent and usable. There is a shift happening where the ability to train a base model to be a general agentic reasoner is becoming opaque like at-scale pretraining practices from a few years ago. This could go so far as to change the established pretraining, midtraining, post-training lexicon that has been standard for a few years. It could come to be something closer to pretraining, reasoning training, and post-training.

随着对训练整个模型的兴趣减少,对投资开源 AI 的兴趣也在减少。这些是我们能得到的唯一线索,但我们无法对抗这些情况的经济引力。这一趋势是发布基础模型(核心推理训练之前的模型版本)的开源模型构建者数量持续减少的下一步。与此同时,开源模型构建者正在尝试对下游产品或推理使用采用收入分成许可。这些实验旨在维持构建接近前沿的开源权重模型的资金可行性——近期很大程度上取决于它们的成功程度。这些人是 Nvidia 围绕开源的需求增长战略成功所必需的,也是最后的关键。

As there is less interest in training the entire model, there is less interest in investing in open-source AI. These are the only sort of hints we will get, but we cannot do much to fight the economic gravity of these situations. This trend is the next step in the number of open model builders who release base models (the model versions before core reasoning training) continuing to decrease. It goes hand in hand with open model builders experimenting with revenue-share licenses for downstream use in products or inference. These are experiments in keeping the financing viable for building near-frontier open-weight models – a lot hinges in the near future on how successful they are. These are the people that need to succeed for Nvidia's demand-growth strategy around open-source to succeed, and last.

在此过程中,开源权重模型领域仍将充满大量活动,因为开放智能访问是最强大的商业策略之一。这种额外类型的参与者,通过间接方式将 AI 变现,以 Meta 和其他拥有庞大资产负债表的超大规模企业为代表。Meta 将其非常强大的 Muse Spark 1.2 模型以开源权重形式发布,将严重阻碍其竞争对手 Anthropic 和 OpenAI 的收入增长率,后者依赖出售 token。这些公司都在将自己的互补品商品化,但方式不同。Nvidia 希望教会每个人自己“钓鱼”获取 token,以使生态系统自给自足,而 Meta 则在战略性地用 token 淹没整个领域。

Along the way we're still in for a ton of action in open-weight models, as releasing access to intelligence is one of the strongest business strategies available. This additional type of player, who monetizes the AI indirectly, is typified by Meta and other hyperscalers with massive balance sheets. Meta releasing its very-strong Muse Spark 1.2 model as open-weights would severely hamper the revenue growth rate of their competitors in Anthropic and OpenAI who rely on selling tokens. These companies are both commoditizing their complements, but they're doing it in different ways. Nvidia wants to teach everyone to fish for tokens, so the ecosystem is self-sustaining, but Meta is strategically flooding the zone with tokens.

互动版:图/公式 + 针对本篇提问 →