2025:大语言模型之年

2025: The year in LLMs

西蒙·威利森 Simon Willison · · 2025-12-31 · Simon Willison's Weblog ↗

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

由 Atlassian 赞助——给你的智能体一个计划,而不是一条提示。新的 Jira 功能为 AI 原生软件开发解锁了完整上下文。现在可以直接从 Jira 将任务分配给 Claude、Cursor 或 GitHub Copilot。了解更多。

Sponsored by: Atlassian — Give your agents a plan. Not a prompt. New Jira capabilities unlock full-context for AI-native software development. Assign tasks to Claude, Cursor, or GitHub Copilot, now directly from Jira. Learn more

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 31)

全文 · Full text(逐段中英对照)

Simon Willison 的博客 Simon Willison’s Weblog(https://simonwillison.net/)

赞助商:Atlassian — 给你的智能体一个计划,而不是一个提示。新的 Jira 功能为 AI 原生软件开发解锁了完整上下文。现在可以直接从 Jira 将任务分配给 Claude、Cursor 或 GitHub Copilot。了解更多

Sponsored by: Atlassian — Give your agents a plan. Not a prompt. New Jira capabilities unlock full-context for AI-native software development. Assign tasks to Claude, Cursor, or GitHub Copilot, now directly from Jira. Learn more

2025:大语言模型年度回顾 2025: The year in LLMs

这是我年度系列文章的第三篇,回顾过去 12 个月 LLM 领域发生的所有大事。前几年的内容请参见《2023 年我们搞清楚的 AI 问题》和《2024 年我们学到的 LLM 知识》。

This is the third in my annual series reviewing everything that happened in the LLM space over the past 12 months. For previous years see Stuff we figured out about AI in 2023 and Things we learned about LLMs in 2024.

这一年充满了各种各样的趋势。

It’s been a year filled with a _lot_ of different trends.

* 编程智能体与 Claude Code 之年

* The year of coding agents and Claude Code

* YOLO 与偏差常态化之年

* The year of YOLO and the Normalization of Deviance

* 中国顶级开源权重模型之年

* The year of top-ranked Chinese open weight models

* 提示驱动图像编辑之年

* The year of prompt-driven image editing

* 模型在学术竞赛中夺金之年

* The year models won gold in academic competitions

* 令人担忧的 AI 赋能浏览器之年

* The year of alarmingly AI-enabled browsers

* 本地模型变好,但云端模型更优之年

* The year local models got good, but cloud models got even better

* 数据中心极度不受欢迎之年

* The year that data centers got extremely unpopular

“推理”之年 The year of “reasoning”

OpenAI 在 2024 年 9 月以 o1 和 o1-mini 开启了“推理”(即推理时缩放/基于可验证奖励的强化学习,RLVR)革命。在 2025 年初,他们又推出了 o3、o3-mini 和 o4-mini,进一步强化了这一方向。此后,推理几乎成为所有主要 AI 实验室模型的标志性功能。

OpenAI kicked off the “reasoning” aka inference-scaling aka Reinforcement Learning from Verifiable Rewards (RLVR) revolution in September 2024 with o1 and o1-mini. They doubled down on that with o3, o3-mini and o4-mini in the opening months of 2025 and reasoning has since become a signature feature of models from nearly every other major AI lab.

我最喜欢的关于这一技巧重要性的解释来自 Andrej Karpathy:

My favourite explanation of the significance of this trick comes from Andrej Karpathy:

2025 年,每个知名的 AI 实验室都至少发布了一个推理模型。有些实验室发布了混合模型,可以在推理模式或非推理模式下运行。许多 API 模型现在都包含用于增加或减少对给定提示应用推理量的调节旋钮。

Every notable AI lab released at least one reasoning model in 2025. Some labs released hybrids that could be run in reasoning or non-reasoning modes. Many API models now include dials for increasing or decreasing the amount of reasoning applied to a given prompt.

我花了一段时间才理解推理的用途。最初的演示展示了它解决数学逻辑谜题和数 strawberry 中 R 的个数——这两件事我在日常使用模型中并不需要。

It took me a while to understand what reasoning was useful for. Initial demos showed it solving mathematical logic puzzles and counting the Rs in strawberry—two things I didn’t find myself needing in my day-to-day model usage.

事实证明,推理的真正解锁在于驱动工具。能够使用工具的推理模型可以规划多步骤任务,执行它们,并持续“推理结果”,从而更新计划以更好地实现预期目标。

It turned out that the real unlock of reasoning was in driving tools. Reasoning models with access to tools can plan out multi-step tasks, execute on them and continue to _reason about the results_ such that they can update their plans to better achieve the desired goal.

一个显著的成果是,AI 辅助搜索现在确实有效了。以前将搜索引擎与 LLM 连接起来效果存疑,但现在我发现即使是我更复杂的研究问题,也常常能通过 ChatGPT 中的 GPT-5 Thinking 得到回答。

A notable result is that AI assisted search actually works now. Hooking up search engines to LLMs had questionable results before, but now I find even my more complex research questions can often be answered by GPT-5 Thinking in ChatGPT.

推理模型在生成和调试代码方面也表现出色。推理技巧意味着它们可以从一个错误开始,逐步遍历代码库的多个层次,找到根本原因。我发现即使是最棘手的 bug,也能被一个能够读取和执行代码(即使是大型复杂代码库)的优秀推理模型诊断出来。

Reasoning models are also exceptional at producing and debugging code. The reasoning trick means they can start with an error and step through many different layers of the codebase to find the root cause. I’ve found even the gnarliest of bugs can be diagnosed by a good reasoner with the ability to read and execute code against even large and complex codebases.

将推理与工具使用结合起来,你就会得到……

Combine reasoning with tool-use and you get...

智能体之年 The year of agents

年初我预测智能体不会成为现实。整个 2024 年大家都在谈论智能体,但几乎没有成功的例子,更令人困惑的是,每个人使用“智能体”这个词时似乎都基于略有不同的定义。

I started the year making a prediction that agents were not going to happen. Throughout 2024 everyone was talking about agents but there were few to no examples of them working, further confused by the fact that everyone using the term “agent” appeared to be working from a slightly different definition from everyone else.

到了九月,我厌倦了因为缺乏明确定义而回避这个词,决定将其视为一个在循环中运行工具以实现目标的 LLM。这让我能够进行富有成效的讨论,这始终是我对这类术语的目标。

By September I’d got fed up of avoiding the term myself due to the lack of a clear definition and decided to treat them as an LLM that runs tools in a loop to achieve a goal. This unblocked me for having productive conversations about them, always my goal for any piece of terminology like that.

我认为智能体不会出现,因为我觉得轻信问题无法解决,而且用 LLM 取代人类员工的想法仍然像是可笑的科幻小说。

I didn’t think agents would happen because I didn’t think the gullibility problem could be solved, and I thought the idea of replacing human staff members with LLMs was still laughable science fiction.

我的预测对了一半:科幻小说中那种能完成任何指令的魔法电脑助手(如《Her》)并未实现……

I was _half_ right in my prediction: the science fiction version of a magic computer assistant that does anything you ask of (Her)) didn’t materialize...

但如果将智能体定义为通过多步骤工具调用执行有用工作的 LLM 系统,那么智能体已经存在,并且被证明非常有用。

But if you define agents as LLM systems that can perform useful work via tool calls over multiple steps then agents are here and they are proving to be extraordinarily useful.

智能体的两个突破性应用领域是编程和搜索。

The two breakout categories for agents have been for coding and for search.

深度研究模式——让 LLM 收集信息并花 15 分钟以上生成详细报告——在年初很流行,但现在已不再流行,因为 GPT-5 Thinking(以及谷歌的“AI 模式”,比糟糕的“AI 概述”好得多)能在更短时间内产生类似结果。我认为这是一种智能体模式,而且效果很好。

The Deep Research pattern—where you challenge an LLM to gather information and it churns away for 15+ minutes building you a detailed report—was popular in the first half of the year but has fallen out of fashion now that GPT-5 Thinking (and Google’s "AI mode", a significantly better product than their terrible "AI overviews") can produce comparable results in a fraction of the time. I consider this to be an agent pattern, and one that works really well.

“编程智能体”模式则更为重要。

The “coding agents” pattern is a much bigger deal.

编程智能体与 Claude Code 之年 The year of coding agents and Claude Code

2025 年最具影响力的事件发生在二月,即 Claude Code 的悄然发布。

The most impactful event of 2025 happened in February, with the quiet release of Claude Code.

我说悄然,是因为它甚至没有自己的博客文章!Anthropic 将 Claude Code 的发布作为其宣布 Claude 3.7 Sonnet 文章中的第二项内容。

I say quiet because it didn’t even get its own blog post! Anthropic bundled the Claude Code release in as the second item in their post announcing Claude 3.7 Sonnet.

(为什么 Anthropic 从 Claude 3.5 Sonnet 直接跳到 3.7?因为他们在 2024 年 10 月对 Claude 3.5 进行了重大升级,但保留了完全相同的名称,导致开发者社区开始将未命名的 3.5 Sonnet v2 称为 3.6。Anthropic 因未能正确命名新模型而浪费了整个版本号!)

(Why did Anthropic jump from Claude 3.5 Sonnet to 3.7? Because they released a major bump to Claude 3.5 in October 2024 but kept the name exactly the same, causing the developer community to start referring to un-named 3.5 Sonnet v2 as 3.6. Anthropic burned a whole version number by failing to properly name their new model!)

Claude Code 是我所称的编程智能体——即能够编写代码、执行代码、检查结果并进一步迭代的 LLM 系统——中最突出的例子。

Claude Code is the most prominent example of what I call coding agents—LLM systems that can write code, execute that code, inspect the results and then iterate further.

各大实验室在 2025 年都推出了自己的命令行界面编程智能体。

The major labs all put out their own CLI coding agents in 2025

供应商无关的选项包括 GitHub Copilot CLI、Amp、OpenCode、OpenHands CLI 和 Pi。Zed、VS Code 和 Cursor 等集成开发环境也在编程智能体集成方面投入了大量精力。

Vendor-independent options include GitHub Copilot CLI, Amp, OpenCode, OpenHands CLI, and Pi. IDEs such as Zed, VS Code and Cursor invested a lot of effort in coding agent integration as well.

我第一次接触编程智能体模式是在 2023 年初 OpenAI 的 ChatGPT 代码解释器——一个内置于 ChatGPT 的系统,允许其在 Kubernetes 沙箱中运行 Python 代码。

My first exposure to the coding agent pattern was OpenAI’s ChatGPT Code Interpreter in early 2023—a system baked into ChatGPT that allowed it to run Python code in a Kubernetes sandbox.

今年,当 Anthropic 终于在 9 月发布他们的等效产品时,我感到很高兴,尽管最初的名字令人困惑,叫做“使用 Claude 创建和编辑文件”。

I was delighted this year when Anthropic finally released their equivalent in September, albeit under the baffling initial name of “Create and edit files with Claude”.

10 月,他们重新利用了那个容器沙箱基础设施,推出了用于网页的 Claude Code,此后我几乎每天都在使用它。

In October they repurposed that container sandbox infrastructure to launch Claude Code for web, which I’ve been using on an almost daily basis ever since.

用于网页的 Claude Code 是我所称的异步编程智能体——一个你可以提示后忘记的系统,它会自行解决问题,并在完成后提交拉取请求。OpenAI 的“Codex cloud”(上周更名为“Codex web”)于 2025 年 5 月早些时候推出。Gemini 在这一类别中的产品叫做 Jules,也于 5 月推出。

Claude Code for web is what I call an asynchronous coding agent—a system you can prompt and forget, and it will work away on the problem and file a Pull Request once it’s done. OpenAI “Codex cloud” (renamed to “Codex web” in the last week) launched earlier in May 2025. Gemini’s entry in this category is called Jules, also launched in May.

我喜欢异步编程智能体这一类别。它们很好地解决了在个人笔记本电脑上运行任意代码执行的安全挑战,而且能够同时启动多个任务——通常从我的手机——并在几分钟后获得不错的结果,这真的很有趣。

I love the asynchronous coding agent category. They’re a great answer to the security challenges of running arbitrary code execution on a personal laptop and it’s really fun being able to fire off multiple tasks at once—often from my phone—and get decent results a few minutes later.

我在《使用 Claude Code 和 Codex 等异步编程智能体进行代码研究项目》和《拥抱并行编程智能体生活方式》中写了更多关于我如何使用这些工具的内容。

I wrote more about how I’m using these in Code research projects with async coding agents like Claude Code and Codex and Embracing the parallel coding agent lifestyle.

命令行上的 LLM 之年 The year of LLMs on the command-line

2024 年,我花了很多时间捣鼓我的 LLM 命令行工具,用于从终端访问 LLM,一直觉得奇怪的是,很少有人认真对待通过 CLI 访问模型——它们与管道等 Unix 机制是如此自然的契合。

In 2024 I spent a lot of time hacking on my LLM command-line tool for accessing LLMs from the terminal, all the time thinking that it was weird that so few people were taking CLI access to models seriously—they felt like such a natural fit for Unix mechanisms like pipes.

也许终端太过古怪和小众,永远无法成为访问 LLM 的主流工具?

Maybe the terminal was just too weird and niche to ever become a mainstream tool for accessing LLMs?

Claude Code 及其同类产品已经明确证明,只要有足够强大的模型和合适的框架,开发者会欣然接受在命令行上使用 LLM。

Claude Code and friends have conclusively demonstrated that developers will embrace LLMs on the command line, given powerful enough models and the right harness.

当 LLM 能直接吐出正确的命令时,像 sed、ffmpeg 和 bash 本身这些语法晦涩的终端命令就不再是入门障碍了,这很有帮助。

It helps that terminal commands with obscure syntax like sed and ffmpeg and bash itself are no longer a barrier to entry when an LLM can spit out the right command for you.

截至 12 月 2 日,Anthropic 声称 Claude Code 的年化收入达到 10 亿美元!我完全没想到一个 CLI 工具能达到接近这样的数字。

As-of December 2nd Anthropic credit Claude Code with $1bn in run-rate revenue! I did _not_ expect a CLI tool to reach anything close to those numbers.

事后看来,也许我当初应该把 LLM 从一个副业提升为关键重点!

With hindsight, maybe I should have promoted LLM from a side-project to a key focus!

YOLO 之年与偏差常态化 The year of YOLO and the Normalization of Deviance

大多数编码智能体的默认设置是要求用户几乎对它们采取的每一个行动进行确认。在一个智能体错误可能删除你的主文件夹或恶意提示注入攻击可能窃取你的凭据的世界里,这种默认设置完全合理。

The default setting for most coding agents is to ask the user for confirmation for almost _every action they take_. In a world where an agent mistake could wipe your home folder or a malicious prompt injection attack could steal your credentials this default makes total sense.

任何尝试过以自动确认模式(即 YOLO 模式——Codex CLI 甚至将 --dangerously-bypass-approvals-and-sandbox 别名为 --yolo)运行其智能体的人都体验过这种权衡:在没有安全防护的情况下使用智能体感觉像是一个完全不同的产品。

Anyone who’s tried running their agent with automatic confirmation (aka YOLO mode—Codex CLI even aliases --dangerously-bypass-approvals-and-sandbox to --yolo) has experienced the trade-off: using an agent without the safety wheels feels like a completely different product.

像 Claude Code for web 和 Codex Cloud 这样的异步编码智能体的一大好处是,它们可以默认以 YOLO 模式运行,因为没有个人电脑需要保护。

A big benefit of asynchronous coding agents like Claude Code for web and Codex Cloud is that they can run in YOLO mode by default, since there’s no personal computer to damage.

我一直以 YOLO 模式运行,尽管深知其中的风险。到目前为止,它还没有让我吃亏……

I run in YOLO mode all the time, despite being _deeply_ aware of the risks involved. It hasn’t burned me yet...

今年我最喜欢的关于大语言模型安全的文章之一是安全研究员 Johann Rehberger 的《AI 中的偏差常态化》。

One of my favourite pieces on LLM security this year is The Normalization of Deviance in AI by security researcher Johann Rehberger.

Johann 描述了“偏差常态化”现象,即反复接触风险行为而没有负面后果,导致个人和组织接受该风险行为为正常。

Johann describes the “Normalization of Deviance” phenomenon, where repeated exposure to risky behaviour without negative consequences leads people and organizations to accept that risky behaviour as normal.

这一概念最初由社会学家 Diane Vaughan 在其研究 1986 年挑战者号航天飞机灾难的工作中提出,该灾难由一个有缺陷的 O 型环引起,而工程师们多年来一直知道这个问题。多次成功的发射使 NASA 的文化不再认真对待这一风险。

This was originally described by sociologist Diane Vaughan as part of her work to understand the 1986 Space Shuttle Challenger disaster, caused by a faulty O-ring that engineers had known about for years. Plenty of successful launches led NASA culture to stop taking that risk seriously.

Johann 认为,我们越长时间以本质上不安全的方式运行这些系统而侥幸无事,我们就越接近我们自己的挑战者号灾难。

Johann argues that the longer we get away with running these systems in fundamentally insecure ways, the closer we are getting to a Challenger disaster of our own.

每月 200 美元订阅的一年 The year of $200/month subscriptions

ChatGPT Plus 最初每月 20 美元的价格,实际上是 Nick Turley 基于 Discord 上的一份 Google 表单投票而仓促决定的。自那以后,这个价格点一直保持不变。

ChatGPT Plus’s original $20/month price turned out to be a snap decision by Nick Turley based on a Google Form poll on Discord. That price point has stuck firmly ever since.

今年出现了一个新的定价先例:Claude Pro Max 20x 计划,每月 200 美元。

This year a new pricing precedent has emerged: the Claude Pro Max 20x plan, at $200/month.

OpenAI 有一个类似的 200 美元计划,名为 ChatGPT Pro。Gemini 则有 Google AI Ultra,每月 249 美元,但首三个月有 124.99 美元/月的折扣。

OpenAI have a similar $200 plan called ChatGPT Pro. Gemini have Google AI Ultra at $249/month with a $124.99/month 3-month starting discount.

这些计划似乎带来了可观的收入,尽管没有任何实验室公布了按层级划分的用户数据。

These plans appear to be driving some serious revenue, though none of the labs have shared figures that break down their subscribers by tier.

我个人曾为 Claude 支付过每月 100 美元,并且一旦我当前的免费额度(来自预览他们的某个模型——谢谢,Anthropic)用完后,就会升级到每月 200 美元的计划。我也从许多其他人那里听说,他们很乐意支付这些价格。

I’ve personally paid $100/month for Claude in the past and will upgrade to the $200/month plan once my current batch of free allowance (from previewing one of their models—thanks, Anthropic) runs out. I’ve heard from plenty of other people who are happy to pay these prices too.

你必须大量使用模型才能花掉 200 美元的 API 额度,因此你可能会认为按 token 付费对大多数人来说更经济。但事实证明,像 Claude Code 和 Codex CLI 这样的工具,一旦你给它们设置更具挑战性的任务,就会消耗大量 token,以至于每月 200 美元的计划提供了相当大的折扣。

You have to use models _a lot_ in order to spend $200 of API credits, so you would think it would make economic sense for most people to pay by the token instead. It turns out tools like Claude Code and Codex CLI can burn through enormous amounts of tokens once you start setting them more challenging tasks, to the point that $200/month offers a substantial discount.

中国顶级开源权重模型之年 The year of top-ranked Chinese open weight models

2024 年,中国 AI 实验室初现生机,主要体现在 Qwen 2.5 和早期的 DeepSeek 上。它们是不错的模型,但并未让人感觉是世界顶尖的。

2024 saw some early signs of life from the Chinese AI labs mainly in the form of Qwen 2.5 and early DeepSeek. They were neat models but didn’t feel world-beating.

这种情况在 2025 年发生了巨大变化。我的“ai-in-china”标签下仅 2025 年就有 67 篇帖子,而且我还错过了年底的一些关键发布(尤其是 GLM-4.7 和 MiniMax-M2.1)。

This changed dramatically in 2025. My ai-in-china tag has 67 posts from 2025 alone, and I missed a bunch of key releases towards the end of the year (GLM-4.7 and MiniMax-M2.1 in particular.)

以下是截至 2025 年 12 月 30 日 Artificial Analysis 对开源权重模型的排名:

Here’s the Artificial Analysis ranking for open weight models as-of 30th December 2025:

GLM-4.7、Kimi K2 Thinking、MiMo-V2-Flash、DeepSeek V3.2、MiniMax-M2.1 均为中国开源权重模型。该榜单中排名最高的非中国模型是 OpenAI 的 gpt-oss-120B(高),位列第六。

GLM-4.7, Kimi K2 Thinking, MiMo-V2-Flash, DeepSeek V3.2, MiniMax-M2.1 are all Chinese open weight models. The highest non-Chinese model in that chart is OpenAI’s gpt-oss-120B (high), which comes in sixth place.

中国模型革命真正始于 2024 年圣诞节,当时 DeepSeek 3 发布,据称训练成本约为 550 万美元。DeepSeek 随后于 1 月 20 日发布了 DeepSeek R1,迅速引发了一场重大的 AI/半导体抛售:英伟达市值蒸发约 5930 亿美元,投资者恐慌地认为 AI 可能不再是美国的垄断领域。

The Chinese model revolution really kicked off on Christmas day 2024 with the release of DeepSeek 3, supposedly trained for around $5.5m. DeepSeek followed that on 20th January with DeepSeek R1 which promptly triggered a major AI/semiconductor selloff: NVIDIA lost ~$593bn in market cap as investors panicked that AI maybe wasn’t an American monopoly after all.

恐慌并未持续——英伟达迅速反弹,如今股价已远高于 DeepSeek R1 发布前的水平。但这仍是一个引人注目的时刻。谁能想到一个开源权重模型的发布会产生如此大的影响?

The panic didn’t last—NVIDIA quickly recovered and today are up significantly from their pre-DeepSeek R1 levels. It was still a remarkable moment. Who knew an open weight model release could have that kind of impact?

DeepSeek 很快迎来了一批令人印象深刻的中国 AI 实验室。我尤其关注了以下这些:

DeepSeek were quickly joined by an impressive roster of Chinese AI labs. I’ve been paying attention to these ones in particular:

这些模型大多不仅是开源权重,而且完全开源,采用 OSI 批准的许可证:Qwen 对其大部分模型使用 Apache 2.0,DeepSeek 和 Z.ai 使用 MIT。

Most of these models aren’t just open weight, they are fully open source under OSI-approved licenses: Qwen use Apache 2.0 for most of their models, DeepSeek and Z.ai use MIT.

其中一些模型与 Claude 4 Sonnet 和 GPT-5 不相上下!

Some of them are competitive with Claude 4 Sonnet and GPT-5!

遗憾的是,没有一家中国实验室发布了完整的训练数据或用于训练模型的代码,但它们发布了详细的研究论文,推动了技术前沿的发展,尤其是在高效训练和推理方面。

Sadly none of the Chinese labs have released their full training data or the code they used to train their models, but they have been putting out detailed research papers that have helped push forward the state of the art, especially when it comes to efficient training and inference.

长任务之年 The year of long tasks

关于 LLM 最有趣的最新图表之一是来自 METR 的“不同 LLM 在 50%时间内能完成的软件工程任务的时间跨度”:

One of the most interesting recent charts about LLMs is Time-horizon of software engineering tasks different LLMscan complete 50% of the time from METR:

该图表展示了人类需要最多 5 小时完成的任务,并绘制了能够独立实现相同目标的模型的演变。如你所见,2025 年在这方面取得了巨大飞跃,GPT-5、GPT-5.1 Codex Max 和 Claude Opus 4.5 能够执行人类需要数小时才能完成的任务——而 2024 年最好的模型最多只能处理 30 分钟以内的任务。

The chart shows tasks that take humans up to 5 hours, and plots the evolution of models that can achieve the same goals working independently. As you can see, 2025 saw some enormous leaps forward here with GPT-5, GPT-5.1 Codex Max and Claude Opus 4.5 able to perform tasks that take humans multiple hours—2024’s best models tapped out at under 30 minutes.

METR 得出结论:“AI 能完成的任务长度每 7 个月翻一番”。我不确定这种模式是否会持续,但这是一种引人注目的方式,用以说明当前智能体能力的趋势。

METR conclude that “the length of tasks AI can do is doubling every 7 months”. I’m not convinced that pattern will continue to hold, but it’s an eye-catching way of illustrating current trends in agent capabilities.

提示驱动图像编辑之年 The year of prompt-driven image editing

有史以来最成功的消费产品发布发生在三月,而该产品甚至没有名字。

The most successful consumer product launch of all time happened in March, and the product didn’t even have a name.

2024 年 5 月 GPT-4o 的标志性特性之一是其多模态输出——“o”代表“omni”,OpenAI 的发布公告中包含了众多“即将推出”的功能,其中模型除了文本还能输出图像。

One of the signature features of GPT-4o in May 2024 was meant to be its multimodal output—the “o” stood for “omni” and OpenAI’s launch announcement included numerous “coming soon” features where the model output images in addition to text.

然后……什么都没有。图像输出功能未能实现。

Then... nothing. The image output feature failed to materialize.

三月我们终于看到了它的能力——尽管形式更像现有的 DALL-E。OpenAI 在 ChatGPT 中提供了这种新的图像生成功能,关键特性是你可以上传自己的图像,并用提示词告诉它如何修改。

In March we finally got to see what this could do—albeit in a shape that felt more like the existing DALL-E. OpenAI made this new image generation available in ChatGPT with the key feature that you could upload your own images and use prompts to tell it how to modify them.

这一新功能在一周内带来了 1 亿 ChatGPT 注册用户。高峰期他们在一小时内看到了 100 万个账户创建!

This new feature was responsible for 100 million ChatGPT signups in a week. At peak they saw 1 million account creations in a single hour!

像“吉卜力化”——将照片修改成宫崎骏电影画面——这样的技巧一次又一次地走红。

Tricks like “ghiblification”—modifying a photo to look like a frame from a Studio Ghibli movie—went viral time and time again.

OpenAI 发布了该模型的 API 版本,名为“gpt-image-1”,随后在十月推出了更便宜的 gpt-image-1-mini,并在 12 月 16 日推出了大幅改进的 gpt-image-1.5。

OpenAI released an API version of the model called “gpt-image-1”, later joined by a cheaper gpt-image-1-mini in October and a much improved gpt-image-1.5 on December 16th.

最值得注意的开源权重竞争对手来自 Qwen,他们在 8 月 4 日发布了 Qwen-Image 生成模型,随后在 8 月 19 日发布了 Qwen-Image-Edit。这个模型可以在(配置良好的)消费级硬件上运行!他们随后在十一月推出了 Qwen-Image-Edit-2511,12 月 30 日推出了 Qwen-Image-2512,这两个我还没试过。

The most notable open weight competitor to this came from Qwen with their Qwen-Image generation model on August 4th followed by Qwen-Image-Edit on August 19th. This one can run on (well equipped) consumer hardware! They followed with Qwen-Image-Edit-2511 in November and Qwen-Image-2512 on 30th December, neither of which I’ve tried yet.

图像生成领域更大的新闻来自 Google,他们推出了 Nano Banana 模型,可通过 Gemini 使用。

The even bigger news in image generation came from Google with their Nano Banana models, available via Gemini.

Google 在三月以“Gemini 2.0 Flash 原生图像生成”的名称预览了早期版本。真正好的版本在 8 月 26 日发布,他们开始谨慎地在公开场合使用代号“Nano Banana”(API 模型称为“Gemini 2.5 Flash Image”)。

Google previewed an early version of this in March under the name “Gemini 2.0 Flash native image generation”. The really good one landed on August 26th, where they started cautiously embracing the codename "Nano Banana" in public (the API model was called "Gemini 2.5 Flash Image").

Nano Banana 引起了人们的注意,因为它能生成有用的文本!它显然也是遵循图像编辑指令最好的模型。

Nano Banana caught people’s attention because _it could generate useful text_! It was also clearly the best model at following image editing instructions.

十一月,Google 完全接受了“Nano Banana”这个名字,发布了 Nano Banana Pro。这个模型不仅能生成文本,还能输出真正有用的详细信息图和其他文本及信息密集的图像。它现在是一个专业级工具。

In November Google fully embraced the “Nano Banana” name with the release of Nano Banana Pro. This one doesn’t just generate text, it can output genuinely useful detailed infographics and other text and information-heavy images. It’s now a professional-grade tool.

Max Woolf 发布了最全面的 Nano Banana 提示指南,随后在十二月发布了 Nano Banana Pro 的必备指南。

Max Woolf published the most comprehensive guide to Nano Banana prompting, and followed that up with an essential guide to Nano Banana Pro in December.

我主要用它给我的照片添加鸮鹦鹉。

I’ve mainly been using it to add kākāpō parrots to my photos.

鉴于这些图像工具非常受欢迎,Anthropic 没有发布或集成类似功能到 Claude 中有点令人惊讶。我认为这进一步证明了他们专注于面向专业工作的 AI 工具,但 Nano Banana Pro 正在迅速证明它对任何工作涉及创建演示文稿或其他视觉材料的人都有价值。

Given how incredibly popular these image tools are it’s a little surprising that Anthropic haven’t released or integrated anything similar into Claude. I see this as further evidence that they’re focused on AI tools for professional work, but Nano Banana Pro is rapidly proving itself to be of value to anyone who’s work involves creating presentations or other visual materials.

模型在学术竞赛中夺金之年 The year models won gold in academic competitions

7 月,OpenAI 和 Google Gemini 的推理模型在国际数学奥林匹克竞赛(IMO)中均获得金牌,这是一项自 1959 年以来(1980 年除外)每年举办的著名数学竞赛。

In July reasoning models from both OpenAI and Google Gemini achieved gold medal performance in the International Math Olympiad, a prestigious mathematical competition held annually (bar 1980) since 1959.

这值得注意,因为 IMO 的题目是专门为该竞赛设计的。这些题目绝不可能已存在于训练数据中!

This was notable because the IMO poses challenges that are designed specifically for that competition. There’s no chance any of these were already in the training data!

同样值得注意的是,这两个模型都没有使用工具——它们的解答完全基于内部知识和基于 token 的推理能力。

It’s also notable because neither of the models had access to tools—their solutions were generated purely from their internal knowledge and token-based reasoning capabilities.

事实证明,足够先进的大语言模型确实能解决数学问题!

Turns out sufficiently advanced LLMs can do math after all!

9 月,OpenAI 和 Gemini 在国际大学生程序设计竞赛(ICPC)中取得了类似成就——同样因为题目新颖、未曾公开而引人注目。这次模型可以访问代码执行环境,但无法访问互联网。

In September OpenAI and Gemini pulled off a similar feat for the International Collegiate Programming Contest (ICPC)—again notable for having novel, previously unpublished problems. This time the models had access to a code execution environment but otherwise no internet access.

我不确定这些竞赛中使用的具体模型是否已公开发布,但 Gemini 的 Deep Think 和 OpenAI 的 GPT-5 Pro 应能提供接近的近似性能。

I don’t believe the exact models used for these competitions have been released publicly, but Gemini’s Deep Think and OpenAI’s GPT-5 Pro should provide close approximations.

Llama 迷失方向的一年 The year that Llama lost its way

事后看来,2024 年是 Llama 的一年。Meta 的 Llama 模型是迄今为止最受欢迎的开源权重模型——最初的 Llama 在 2023 年开启了开源权重革命,而 Llama 3 系列,尤其是 3.1 和 3.2 小版本,在开源权重能力上实现了巨大飞跃。

With hindsight, 2024 was the year of Llama. Meta’s Llama models were by far the most popular open weight models—the original Llama kicked off the open weight revolution back in 2023 and the Llama 3 series, in particular the 3.1 and 3.2 dot-releases, were huge leaps forward in open weight capability.

Llama 4 被寄予厚望,但当它在 4 月发布时……却有些令人失望。

Llama 4 had high expectations, and when it landed in April it was... kind of disappointing.

有一个小丑闻:在 LMArena 上测试的模型并非实际发布的模型,但我的主要抱怨是模型_太大了_。之前 Llama 版本最棒的一点是它们通常包含可以在笔记本电脑上运行的尺寸。Llama 4 Scout 和 Maverick 模型分别为 109B 和 400B,大到即使量化也无法在我的 64GB Mac 上运行。

There was a minor scandal where the model tested on LMArena turned out not to be the model that was released, but my main complaint was that the models were _too big_. The neatest thing about previous Llama releases was that they often included sizes you could run on a laptop. The Llama 4 Scout and Maverick models were 109B and 400B, so big that even quantization wouldn’t get them running on my 64GB Mac.

它们是用 2T Llama 4 Behemoth 训练的,这个模型似乎已被遗忘——它当然没有被发布。

They were trained using the 2T Llama 4 Behemoth which seems to have been forgotten now—it certainly wasn’t released.

LM Studio 列出的最受欢迎模型中没有一个是来自 Meta 的,这很能说明问题;而 Ollama 上最受欢迎的仍然是 Llama 3.1,它在排行榜上的位置也很低。

It says a lot that none of the most popular models listed by LM Studio are from Meta, and the most popular on Ollama is still Llama 3.1, which is low on the charts there too.

Meta 今年的 AI 新闻主要涉及内部政治和花费巨额资金为其新的超级智能实验室招聘人才。目前尚不清楚未来是否会有 Llama 版本发布,或者他们是否已从开源权重模型发布转向专注于其他事情。

Meta’s AI news this year mainly involved internal politics and vast amounts of money spent hiring talent for their new Superintelligence Labs. It’s not clear if there are any future Llama releases in the pipeline or if they’ve moved away from open weight model releases to focus on other things.

OpenAI 失去领先地位的一年 The year that OpenAI lost their lead

去年,OpenAI 在大型语言模型领域仍然是无可争议的领导者,尤其是考虑到 o1 及其 o3 推理模型的预览。

Last year OpenAI remained the undisputed leader in LLMs, especially given o1 and the preview of their o3 reasoning models.

今年,行业其他公司迎头赶上。

This year the rest of the industry caught up.

OpenAI 仍然拥有顶级模型,但它们正面临全方位的挑战。

OpenAI still have top tier models, but they’re being challenged across the board.

在图像模型方面,它们仍然被 Nano Banana Pro 击败。在代码方面,许多开发者认为 Opus 4.5 略优于 GPT-5.2 Codex。在开放权重模型方面,它们的 gpt-oss 模型虽然出色,但正落后于中国的人工智能实验室。它们在音频领域的领先地位正受到 Gemini Live API 的威胁。

In image models they’re still being beaten by Nano Banana Pro. For code a lot of developers rate Opus 4.5 very slightly ahead of GPT-5.2 Codex. In open weight models their gpt-oss models, while great, are falling behind the Chinese AI labs. Their lead in audio is under threat from the Gemini Live API.

OpenAI 的胜出之处在于消费者的心智份额。没有人知道什么是“LLM”,但几乎每个人都听说过 ChatGPT。它们的消费者应用在用户数量上仍然远超 Gemini 和 Claude。

Where OpenAI are winning is in consumer mindshare. Nobody knows what an “LLM” is but almost everyone has heard of ChatGPT. Their consumer apps still dwarf Gemini and Claude in terms of user numbers.

它们在此面临的最大风险是 Gemini。去年 12 月,OpenAI 宣布进入红色警戒状态以应对 Gemini 3,推迟了新项目的开发,专注于与关键产品的竞争。

Their biggest risk here is Gemini. In December OpenAI declared a Code Red in response to Gemini 3, delaying work on new initiatives to focus on the competition with their key products.

双子座之年 The year of Gemini

他们在此发布了胜利的 2025 年回顾。2025 年见证了 Gemini 2.0、Gemini 2.5 以及 Gemini 3.0——每个模型系列都支持超过 100 万个 token 的音频/视频/图像/文本输入,定价具有竞争力,并且能力逐代增强。

They posted their own victorious 2025 recap here. 2025 saw Gemini 2.0, Gemini 2.5 and then Gemini 3.0—each model family supporting audio/video/image/text input of 1,000,000+ tokens, priced competitively and proving more capable than the last.

他们还推出了 Gemini CLI(他们的开源命令行编码智能体,已被 Qwen 用于 Qwen Code 的分支)、Jules(他们的异步编码智能体)、AI Studio 的持续改进、Nano Banana 图像模型、用于视频生成的 Veo 3、有前景的 Gemma 3 开放权重模型系列以及一系列较小的功能。

They also shipped Gemini CLI (their open source command-line coding agent, since forked by Qwen for Qwen Code), Jules (their asynchronous coding agent), constant improvements to AI Studio, the Nano Banana image models, Veo 3 for video generation, the promising Gemma 3 family of open weight models and a stream of smaller features.

谷歌最大的优势在于幕后。几乎所有其他 AI 实验室都使用 NVIDIA GPU 进行训练,这些 GPU 以支撑 NVIDIA 数万亿美元估值的利润率出售。

Google’s biggest advantage lies under the hood. Almost every other AI lab trains with NVIDIA GPUs, which are sold at a margin that props up NVIDIA’s multi-trillion dollar valuation.

谷歌使用自己的内部硬件 TPU,今年他们证明了这些 TPU 在模型的训练和推理方面都表现出色。

Google use their own in-house hardware, TPUs, which they’ve demonstrated this year work exceptionally well for both training and inference of their models.

当你最大的开销是花在 GPU 上的时间时,拥有一个拥有自己优化且可能更便宜的硬件堆栈的竞争对手是一个令人生畏的前景。

When your number one expense is time spent on GPUs, having a competitor with their own, optimized and presumably much cheaper hardware stack is a daunting prospect.

让我觉得有趣的是,Google Gemini 是产品名称反映公司内部组织架构的终极例子——它被称为 Gemini,是因为它源于谷歌 DeepMind 和谷歌大脑团队的合并(如同双胞胎)。

It continues to tickle me that Google Gemini is the ultimate example of a product name that reflects the company’s internal org-chart—it’s called Gemini because it came out of the bringing together (as twins) of Google’s DeepMind and Google Brain teams.

鹈鹕骑自行车之年 The year of pelicans riding bicycles

我在 2024 年 10 月首次要求一个 LLM 生成一只鹈鹕骑自行车的 SVG,但 2025 年才是真正深入的时候。它最终成了一个梗。

I first asked an LLM to generate an SVG of a pelican riding a bicycle in October 2024, but 2025 is when I really leaned into it. It’s ended up a meme in its own right.

我最初只是想开个愚蠢的玩笑。自行车很难画,鹈鹕也是,而且鹈鹕的形状不适合骑自行车。我相当确定训练数据中不会有相关的内容,所以让一个文本输出模型生成一个 SVG 插图感觉是一个相当荒谬的挑战。

I originally intended it as a dumb joke. Bicycles are hard to draw, as are pelicans, and pelicans are the wrong shape to ride a bicycle. I was pretty sure there wouldn’t be anything relevant in the training data, so asking a text-output model to generate an SVG illustration of one felt like a somewhat absurdly difficult challenge.

令我惊讶的是,模型画鹈鹕骑自行车的能力与其整体能力之间似乎存在相关性。

To my surprise, there appears to be a correlation between how good the model is at drawing pelicans on bicycles and how good it is overall.

我对此并没有真正的解释。这个模式是在我为 7 月份的 AI 工程师世界博览会准备一个临时的主题演讲(因为有演讲者退出)时才变得清晰的。

I don’t really have an explanation for this. The pattern only became clear to me when I was putting together a last-minute keynote (they had a speaker drop out) for the AI Engineer World’s Fair in July.

你可以在这里阅读(或观看)我的演讲:过去六个月的大语言模型,以鹈鹕骑自行车为例。

You can read (or watch) the talk I gave here: The last six months in LLMs, illustrated by pelicans on bicycles.

我的完整插图集可以在我的“鹈鹕骑自行车”标签下找到——已有 89 篇文章,并且还在增加。

My full collection of illustrations can be found on my pelican-riding-a-bicycle tag—89 posts and counting.

有大量证据表明 AI 实验室知道这个基准。它在 5 月份的 Google I/O 主题演讲中短暂出现,在 10 月份的 Anthropic 可解释性研究论文中被提及,8 月份我还在 OpenAI 总部拍摄的 GPT-5 发布视频中谈到了它。

There is plenty of evidence that the AI labs are aware of the benchmark. It showed up (for a split second) in the Google I/O keynote in May, got a mention in an Anthropic interpretability research paper in October and I got to talk about it in a GPT-5 launch video filmed at OpenAI HQ in August.

他们是否专门针对这个基准进行训练?我不这么认为,因为即使是最先进的前沿模型生成的鹈鹕插图仍然很糟糕!

Are they training specifically for the benchmark? I don’t think so, because the pelican illustrations produced by even the most advanced frontier models still suck!

在《如果 AI 实验室训练鹈鹕骑自行车会发生什么?》中,我坦白了我的险恶目的:

In What happens if AI labs train for pelicans riding bicycles? I confessed to my devious objective:

我最喜欢的仍然是这张来自 GPT-5 的:

My favourite is still this one that I go from GPT-5:

我构建了 110 个工具的那一年 The year I built 110 tools

我去年创建了 tools.simonwillison.net 网站,作为我不断增长的 vibe-coding / AI 辅助 HTML+JavaScript 工具集的单一存放位置。我在这一年中写了几篇较长的文章:

I started my tools.simonwillison.net site last year as a single location for my growing collection of vibe-coded / AI-assisted HTML+JavaScript tools. I wrote several longer pieces about this throughout the year:

* 这里是我如何使用 LLM 帮助我编写代码

* Here’s how I use LLMs to help me write code

* 为我的工具集添加 AI 生成的描述

* Adding AI-generated descriptions to my tools collection

* 使用 Claude Code for web 构建一个复制粘贴分享终端会话的工具

* Building a tool to copy-paste share terminal sessions using Claude Code for web

* 构建 HTML 工具的有用模式——这是我这一系列中最喜欢的一篇。

* Useful patterns for building HTML tools—my favourite post of the bunch.

新的按月浏览页面显示,我在 2025 年构建了 110 个这样的工具!

The new browse all by month page shows I built 110 of these in 2025!

我非常喜欢这种构建方式,并且认为这是练习和探索这些模型能力的绝佳方式。几乎每个工具都附带了提交历史,链接到我用于构建它们的提示词和对话记录。

I really enjoy building in this way, and I think it’s a fantastic way to practice and explore the capabilities of these models. Almost every tool is accompanied by a commit history that links to the prompts and transcripts I used to build them.

我将重点介绍过去一年中我最喜欢的一些工具:

I’ll highlight a few of my favourites from the past year:

* blackened-cauliflower-and-turkish-style-stew 很荒谬。它是一个自定义烹饪计时器应用,适用于需要同时准备 Green Chef 的 Blackened Cauliflower 和 Turkish-style Spiced Chickpea Stew 食谱的人。这里有更多关于它的信息。

* blackened-cauliflower-and-turkish-style-stew is ridiculous. It’s a custom cooking timer app for anyone who needs to prepare Green Chef’s Blackened Cauliflower and Turkish-style Spiced Chickpea Stew recipes at the same time. Here’s more about that one.

* is-it-a-bird 灵感来自 xkcd 1425,通过 Transformers.js 加载一个 150MB 的 CLIP 模型,并用它来判断一张图片或网络摄像头画面是否是鸟。

* is-it-a-bird takes inspiration from xkcd 1425, loads a 150MB CLIP model via Transformers.js and uses it to say if an image or webcam feed is a bird or not.

* bluesky-thread 让我可以查看 Bluesky 上的任何帖子串,并带有“最新优先”选项,以便更容易地跟踪新发布的帖子。

* bluesky-thread lets me view any thread on Bluesky with a “most recent first” option to make it easier to follow new posts as they arrive.

其他许多工具对我的工作流程很有用,比如 svg-render、render-markdown 和 alt-text-extractor。我还构建了一个工具,针对 localStorage 进行隐私友好的个人分析,以跟踪我最常使用哪些工具。

A lot of the others are useful tools for my own workflow like svg-render and render-markdown and alt-text-extractor. I built one that does privacy-friendly personal analytics against localStorage to keep track of which tools I use the most often.

告密之年! The year of the snitch!

Anthropic 为其模型发布的系统卡片一直值得全文阅读——它们包含大量有用信息,并且常常会滑入有趣的科幻领域。

Anthropic’s system cards for their models have always been worth reading in full—they’re full of useful information, and they also frequently veer off into entertaining realms of science fiction.

五月份的 Claude 4 系统卡片中有一些特别有趣的时刻——重点标记是我加的:

The Claude 4 system card in May had some particularly fun moments—highlights mine:

换句话说,Claude 4 可能会向联邦政府告发你。

In other words, Claude 4 might snitch you out to the feds.

这引起了大量媒体关注,许多人谴责 Anthropic 训练了一个过于道德的模型。随后 Theo Browne 利用系统卡片中的概念构建了 SnitchBench——一个用于衡量不同模型向用户告密可能性的基准测试。

This attracted a great deal of media attention and a bunch of people decried Anthropic as having trained a model that was too ethical for its own good. Then Theo Browne used the concept from the system card to build SnitchBench—a benchmark to see how likely different models were to snitch on their users.

结果发现_它们几乎都会做同样的事情_!

It turns out _they almost all do the same thing_!

Theo 制作了一个视频,我也发表了自己用我的 LLM 复现 SnitchBench 的笔记。

Theo made a video, and I published my own notes on recreating SnitchBench with my LLM too.

我建议不要把它放在你的系统提示中!Anthropic 最初的 Claude 4 系统卡片也说了同样的话:

I recommend not putting that in your system prompt! Anthropic’s original Claude 4 system card said the same thing:

氛围编程之年 The year of vibe coding

在二月的一条推文中,Andrej Karpathy 创造了“氛围编程”一词,其定义不幸地冗长(我怀念 140 个字符的时代),许多人未能完整阅读:

In a tweet in February Andrej Karpathy coined the term “vibe coding”, with an unfortunately long definition (I miss the 140 character days) that many people failed to read all the way to the end:

这里的关键思想是“忘记代码的存在”——氛围编程捕捉了一种新的、有趣的软件原型设计方式,仅通过提示即可实现“基本可用”。

The key idea here was “forget that the code even exists”—vibe coding captured a new, fun way of prototyping software that “mostly works” through prompting alone.

我不知道自己是否曾见过一个新术语如此迅速地流行起来——或被扭曲。

I don’t know if I’ve ever seen a new term catch on—or get distorted—so quickly in my life.

许多人反而将氛围编程当作一个笼统的术语,涵盖任何涉及 LLM 的编程。我认为这是对这个好词的浪费,尤其是因为越来越清楚的是,在不久的将来,大多数编程都将涉及某种程度的 AI 辅助。

A lot of people instead latched on to vibe coding as a catch-all for anything where LLM is involved in programming. I think that’s a waste of a great term, especially since it’s becoming clear likely that most programming will involve some level of AI-assistance in the near future.

因为我热衷于与语言风车作战,我尽力鼓励该术语的原始含义:

Because I’m a sucker for tilting at linguistic windmills I tried my best to encourage the original meaning of the term:

* 并非所有 AI 辅助编程都是氛围编程(但氛围编程很棒)——三月

* Not all AI-assisted programming is vibe coding (but vibe coding rocks) in March

* 两家出版社和三位作者在五月未能理解“氛围编程”的含义(其中一本书随后将标题改为更好的“超越氛围编程”)。

* Two publishers and three authors fail to understand what “vibe coding” means in May (one book subsequently changed its title to the much better “Beyond Vibe Coding”).

* 氛围工程——十月,我试图提出一个替代术语,用于描述专业工程师使用 AI 辅助构建生产级软件的情况。

* Vibe engineering in October, where I tried to suggest an alternative term for what happens when professional engineers use AI assistance to build production-grade software.

* 你的工作是交付你已证明能正常工作的代码——十二月,关于专业软件开发是关于可证明正常工作的代码,无论你如何构建它。

* Your job is to deliver code you have proven to work in December, about how professional software development is about code that demonstrably works, no matter how you built it.

我不认为这场战斗已经结束。我看到了令人欣慰的信号,表明更好、更原始的氛围编程定义可能会胜出。

I don’t think this battle is over yet. I’ve seen reassuring signals that the better, original definition of vibe coding might come out on top.

我真的应该找一个不那么对抗性的语言爱好!

I should really get a less confrontational linguistic hobby!

MCP 的(唯一?)一年 The (only?) year of MCP

Anthropic 于 2024 年 11 月发布了他们的模型上下文协议规范,作为将工具调用与不同大语言模型集成的开放标准。2025 年初,它迅速流行起来。5 月时,OpenAI、Anthropic 和 Mistral 在八天内相继推出了对 MCP 的 API 级支持!

Anthropic introduced their Model Context Protocol specification in November 2024 as an open standard for integrating tool calls with different LLMs. In early 2025 it _exploded_ in popularity. There was a point in May where OpenAI, Anthropic, and Mistral all rolled out API-level support for MCP within eight days of each other!

MCP 是一个相当合理的想法,但如此广泛的采用让我感到惊讶。我认为这归结于时机:MCP 的发布恰逢模型在工具调用方面终于变得出色且可靠,以至于许多人似乎误以为 MCP 支持是模型使用工具的先决条件。

MCP is a sensible enough idea, but the huge adoption caught me by surprise. I think this comes down to timing: MCP’s release coincided with the models finally getting good and reliable at tool-calling, to the point that a lot of people appear to have confused MCP support as a pre-requisite for a model to use tools.

有一段时间,MCP 还感觉像是那些面临“AI 战略”压力但不知如何着手的公司的便捷答案。为你的产品宣布一个 MCP 服务器是一种易于理解的方式来满足这一要求。

For a while it also felt like MCP was a convenient answer for companies that were under pressure to have “an AI strategy” but didn’t really know how to do that. Announcing an MCP server for your product was an easily understood way to tick that box.

我认为 MCP 可能只是一年奇迹的原因是编码智能体的迅猛增长。似乎任何情况下最好的工具都是 Bash——如果你的智能体可以运行任意 shell 命令,它就能做任何可以通过在终端中输入命令来完成的事情。

The reason I think MCP may be a one-year wonder is the stratospheric growth of coding agents. It appears that the best possible tool for any situation is Bash—if your agent can run arbitrary shell commands, it can do anything that can be done by typing commands into a terminal.

自从我深入使用 Claude Code 及其同类工具后,我几乎没怎么用过 MCP——我发现像 gh 这样的 CLI 工具和像 Playwright 这样的库是比 GitHub 和 Playwright MCP 更好的替代品。

Since leaning heavily into Claude Code and friends myself I’ve hardly used MCP at all—I’ve found CLI tools like gh and libraries like Playwright to be better alternatives to the GitHub and Playwright MCPs.

Anthropic 自己似乎也在今年晚些时候承认了这一点,他们发布了出色的 Skills 机制——请参阅我十月的文章《Claude Skills 太棒了,也许比 MCP 更重要》。MCP 涉及 Web 服务器和复杂的 JSON 负载。而一个 Skill 只是一个文件夹中的 Markdown 文件,可选地附带一些可执行脚本。

Anthropic themselves appeared to acknowledge this later in the year with their release of the brilliant Skills mechanism—see my October post Claude Skills are awesome, maybe a bigger deal than MCP. MCP involves web servers and complex JSON payloads. A Skill is a Markdown file in a folder, optionally accompanied by some executable scripts.

然后,在 11 月,Anthropic 发布了《使用 MCP 进行代码执行:构建更高效的智能体》——描述了一种让编码智能体生成代码来调用 MCP 的方法,从而避免了原始规范中的许多上下文开销。

Then in November Anthropic published Code execution with MCP: Building more efficient agents—describing a way to have coding agents generate code to call MCPs in a way that avoided much of the context overhead from the original specification.

(我自豪的是,我在 Anthropic 发布 Skills 前一周逆向工程了它们,并在两个月后对 OpenAI 悄然采用 Skills 做了同样的事情。)

(I’m proud of the fact that I reverse-engineered Anthropic’s skills a week before their announcement, and then did the same thing to OpenAI’s quiet adoption of skills two months after that.)

MCP 于 12 月初被捐赠给新的 Agentic AI Foundation。Skills 于 12 月 18 日升级为“开放格式”。

MCP was donated to the new Agentic AI Foundation at the start of December. Skills were promoted to an “open format” on December 18th.

令人警惕的 AI 赋能浏览器之年 The year of alarmingly AI-enabled browsers

尽管存在非常明显的安全风险,但似乎每个人都想把大语言模型放进你的网络浏览器中。

Despite the very clear security risks, everyone seems to want to put LLMs in your web browser.

OpenAI 在 10 月推出了 ChatGPT Atlas,其团队包括长期在 Google Chrome 工作的工程师 Ben Goodger 和 Darin Fisher。

OpenAI launched ChatGPT Atlas in October, built by a team including long-time Google Chrome engineers Ben Goodger and Darin Fisher.

Anthropic 一直在推广他们的 Claude 在 Chrome 扩展中,提供类似的功能作为扩展,而不是完整的 Chrome 分支。

Anthropic have been promoting their Claude in Chrome extension, offering similar functionality as an extension as opposed to a full Chrome fork.

Chrome 本身现在右上角有一个小小的“Gemini”按钮,称为 Gemini in Chrome,尽管我认为它只是用于回答关于内容的问题,目前还没有驱动浏览操作的能力。

Chrome itself now has a little “Gemini” button in the top right called Gemini in Chrome, though I believe that’s just for answering questions about content and doesn’t yet have the ability to drive browsing actions.

我仍然对这些新工具的安全影响深感担忧。我的浏览器可以访问我最敏感的数据,并控制我大部分的数字生活。针对能够窃取或修改数据的浏览智能体的提示注入攻击是一个可怕的前景。

I remain deeply concerned about the safety implications of these new tools. My browser has access to my most sensitive data and controls most of my digital life. A prompt injection attack against a browsing agent that can exfiltrate or modify that data is a terrifying prospect.

到目前为止,我在缓解这些担忧方面看到的最详细的细节来自 OpenAI 的首席信息安全官 Dane Stuckey,他谈到了护栏、红队测试和纵深防御,但也正确地指出提示注入是“一个前沿的、未解决的安全问题”。

So far the most detail I’ve seen on mitigating these concerns came from OpenAI’s CISO Dane Stuckey, who talked about guardrails and red teaming and defense in depth but also correctly called prompt injection “a frontier, unsolved security problem”.

我已经在非常严密的监督下使用过几次这些浏览器智能体(例如)。它们有点慢且不稳定——在点击交互元素时常常失误——但对于解决无法通过 API 处理的问题来说很方便。

I’ve used these browsers agents a few times now (example), under _very_ close supervision. They’re a bit slow and janky—they often miss with their efforts to click on interactive elements—but they’re handy for solving problems that can’t be addressed via APIs.

我仍然对它们感到不安,尤其是在那些不如我多疑的人手中。

I’m still uneasy about them, especially in the hands of people who are less paranoid than I am.

致命三重奏之年 The year of the lethal trifecta

我写关于提示注入攻击的文章已经三年多了。我持续遇到的一个挑战是,帮助人们理解为什么这是任何在这个领域构建软件的人都需要认真对待的问题。

I’ve been writing about prompt injection attacks for more than three years now. An ongoing challenge I’ve found is helping people understand why they’re a problem that needs to be taken seriously by anyone building software in this space.

语义扩散并没有帮助,因为“提示注入”一词已经扩展到涵盖越狱(尽管我反对),而且谁真的在乎有人能诱使模型说些粗鲁的话呢?

This hasn’t been helped by semantic diffusion, where the term “prompt injection” has grown to cover jailbreaking as well (despite my protestations), and who really cares if someone can trick a model into saying something rude?

所以我尝试了一种新的语言技巧!六月我创造了“致命三重奏”这个术语,用来描述提示注入的一个子集,其中恶意指令诱使智能体代表攻击者窃取私人数据。

So I tried a new linguistic trick! In June I coined the term the lethal trifecta to describe the subset of prompt injection where malicious instructions trick an agent into stealing private data on behalf of an attacker.

我在这里使用的技巧是,人们会直接跳到他们听到的任何新术语的最明显定义。“提示注入”听起来像是“注入提示”。“致命三重奏”故意含糊其辞:如果你想了解它的含义,就必须去寻找我的定义!

A trick I use here is that people will jump straight to the most obvious definition of any new term that they hear. “Prompt injection” sounds like it means “injecting prompts”. “The lethal trifecta” is deliberately ambiguous: you have to go searching for my definition if you want to know what it means!

这似乎奏效了。今年我看到不少人在谈论致命三重奏,到目前为止,没有对其意图含义的误解。

It seems to have worked. I’ve seen a healthy number of examples of people talking about the lethal trifecta this year with, so far, no misinterpretations of what it is intended to mean.

在手机上编程的一年 The year of programming on my phone

今年我在手机上写的代码量远超在电脑上写的。

I wrote significantly more code on my phone this year than I did on my computer.

今年大部分时间里,这主要是因为我很投入于“氛围编程”。我的 tools.simonwillison.net 上的 HTML+JavaScript 工具集大多是这样构建的:我有个小项目的想法,通过各自的 iPhone 应用提示 Claude Artifacts 或 ChatGPT(最近还有 Claude Code),然后要么复制结果粘贴到 GitHub 的网页编辑器,要么等待生成一个拉取请求,我可以在 Mobile Safari 中审查并合并。

Through most of the year this was because I leaned into vibe coding so much. My tools.simonwillison.net collection of HTML+JavaScript tools was mostly built this way: I would have an idea for a small project, prompt Claude Artifacts or ChatGPT or (more recently) Claude Code via their respective iPhone apps, then either copy the result and paste it into GitHub’s web editor or wait for a PR to be created that I could then review and merge in Mobile Safari.

那些 HTML 工具通常只有 100-200 行代码,充斥着无趣的样板代码和重复的 CSS 与 JavaScript 模式——但 110 个这样的工具加起来就很多了!

Those HTML tools are often ~100-200 lines of code, full of uninteresting boilerplate and duplicated CSS and JavaScript patterns—but 110 of them adds up to a lot!

直到 11 月,我还会说我在手机上写了更多代码,但我在笔记本电脑上写的代码显然更重要——经过全面审查、测试更充分,且用于生产环境。

Up until November I would have said that I wrote more code on my phone, but the code I wrote on my laptop was clearly more significant—fully reviewed, better tested and intended for production use.

在过去一个月里,我对 Claude Opus 4.5 的信心大增,开始用手机上的 Claude Code 处理更复杂的任务,包括那些我打算用于非玩具项目的代码。

In the past month I’ve grown confident enough in Claude Opus 4.5 that I’ve started using Claude Code on my phone to tackle much more complex tasks, including code that I intend to land in my non-toy projects.

这始于我将 JustHTML HTML5 解析器从 Python 移植到 JavaScript 的项目,使用了 Codex CLI 和 GPT-5.2。当仅通过提示就能成功时,我开始好奇如果用手机完成类似项目能做到什么程度。

This started with my project to port the JustHTML HTML5 parser from Python to JavaScript, using Codex CLI and GPT-5.2. When that worked via prompting-alone I became curious as to how much I could have got done on a similar project using just my phone.

于是我尝试将 Fabrice Bellard 的新 MicroQuickJS C 库移植到 Python,全程使用 iPhone 上的 Claude Code……而且基本成功了!

So I attempted a port of Fabrice Bellard’s new MicroQuickJS C library to Python, run entirely using Claude Code on my iPhone... and it mostly worked!

这是我会用于生产环境的代码吗?对于不可信代码当然还不至于,但我信任它来执行我自己编写的 JavaScript。我从 MicroQuickJS 借用的测试套件给了我一些信心。

Is it code that I’d use in production? Certainly not yet for untrusted code, but I’d trust it to execute JavaScript I’d written myself. The test suite I borrowed from MicroQuickJS gives me some confidence there.

合规套件之年 The year of conformance suites

事实证明这是一个重大突破:如果你能为最新的编码智能体提供一套现有的测试套件,它们针对~2025 年 11 月的前沿模型表现非常出色。我称这些为合规套件,并开始有意寻找它们——到目前为止,我已经在 html5lib 测试、MicroQuickJS 测试套件以及一个尚未发布的项目(针对全面的 WebAssembly 规范/测试集合)上取得了成功。

This turns out to be the big unlock: the latest coding agents against the ~November 2025 frontier models are remarkably effective if you can give them an existing test suite to work against. I call these conformance suites and I’ve started deliberately looking out for them—so far I’ve had success with the html5lib tests, the MicroQuickJS test suite and a not-yet-released project against the comprehensive WebAssembly spec/test collection.

如果你在 2026 年向世界引入新的协议甚至新的编程语言,我强烈建议将语言无关的合规套件作为项目的一部分。

If you’re introducing a new protocol or even a new programming language to the world in 2026 I strongly recommend including a language-agnostic conformance suite as part of your project.

我见过很多担忧,认为需要被包含在 LLM 训练数据中意味着新技术将难以获得采用。我希望合规套件方法能够帮助缓解这个问题,并使这种形式的新想法更容易获得关注。

I’ve seen plenty of hand-wringing that the need to be included in LLM training data means new technologies will struggle to gain adoption. My hope is that the conformance suite approach can help mitigate that problem and make it _easier_ for new ideas of that shape to gain traction.

本地模型变好的一年,但云端模型变得更好 The year local models got good, but cloud models got even better

2024 年底,我对在自己的机器上运行本地大语言模型失去了兴趣。12 月的 Llama 3.3 70B 重新点燃了我的兴趣,这是我第一次觉得可以在我的 64GB MacBook Pro 上运行一个真正 GPT-4 级别的模型。

Towards the end of 2024 I was losing interest in running local LLMs on my own machine. My interest was re-kindled by Llama 3.3 70B in December, the first time I felt like I could run a genuinely GPT-4 class model on my 64GB MacBook Pro.

然后在一月份,Mistral 发布了 Mistral Small 3,这是一个 Apache 2 许可的 24B 参数模型,它似乎用大约三分之一的内存就达到了与 Llama 3.3 70B 相同的效果。现在我可以运行一个~GPT-4 级别的模型,并且还有剩余内存来运行其他应用!

Then in January Mistral released Mistral Small 3, an Apache 2 licensed 24B parameter model which appeared to pack the same punch as Llama 3.3 70B using around a third of the memory. Now I could run a ~GPT-4 class model and have memory left over to run other apps!

这一趋势在整个 2025 年持续,尤其是当中国 AI 实验室的模型开始占据主导地位时。那个~20-32B 参数的甜蜜点不断产生比上一个更好的模型。

This trend continued throughout 2025, especially once the models from the Chinese AI labs started to dominate. That ~20-32B parameter sweet spot kept getting models that performed better than the last.

我离线完成了一些真正的工作!我对本地大语言模型的热情确实重新燃起了。

I got small amounts of real work done offline! My excitement for local LLMs was very much rekindled.

问题是,大型云端模型也变得更好——包括那些开放权重的模型,虽然它们可以免费获取,但规模太大(100B+),无法在我的笔记本电脑上运行。

The problem is that the big cloud models got better too—including those open weight models that, while freely available, were far too large (100B+) to run on my laptop.

编码智能体改变了我的一切。像 Claude Code 这样的系统需要的不仅仅是一个优秀的模型——它们需要一个推理模型,能够在不断扩展的上下文窗口中可靠地执行数十甚至数百次工具调用。

Coding agents changed everything for me. Systems like Claude Code need more than a great model—they need a reasoning model that can perform reliable tool calling invocations dozens if not hundreds of times over a constantly expanding context window.

我还没有尝试过一个本地模型能够足够可靠地处理 Bash 工具调用,以至于我可以信任该模型在我的设备上操作编码智能体。

I have yet to try a local model that handles Bash tool calls reliably enough for me to trust that model to operate a coding agent on my device.

我的下一台笔记本电脑将至少有 128GB 内存,所以 2026 年的某个开放权重模型有可能满足要求。不过目前,我仍然使用最好的前沿托管模型作为我的日常驱动。

My next laptop will have at least 128GB of RAM, so there’s a chance that one of the 2026 open weight models might fit the bill. For now though I’m sticking with the best available frontier hosted models as my daily drivers.

垃圾内容之年 The year of slop

我在 2024 年帮助推广“垃圾内容”一词方面发挥了很小的作用,在 5 月撰写了相关文章,并随后在《卫报》和《纽约时报》上获得了引用。

I played a tiny role helping to popularize the term “slop” in 2024, writing about it in May and landing quotes in the Guardian and the New York Times shortly afterwards.

今年,韦氏词典将其评为年度词汇!

This year Merriam-Webster crowned it word of the year!

我喜欢它代表了一种广泛认同的感觉,即低质量的 AI 生成内容是不好的,应该避免。

I like that it represents a widely understood feeling that poor quality AI-generated content is bad and should be avoided.

我仍然抱有希望,认为垃圾内容最终不会像许多人担心的那样成为一个严重的问题。

I’m still holding hope that slop won’t end up as bad a problem as many people fear.

互联网一直充斥着低质量内容。挑战一如既往,是找到并放大优质内容。我不认为垃圾内容数量的增加会改变这一基本动态。策展比以往任何时候都更重要。

The internet has _always_ been flooded with low quality content. The challenge, as ever, is to find and amplify the good stuff. I don’t see the increased volume of junk as changing that fundamental dynamic much. Curation matters more than ever.

话虽如此……我不使用 Facebook,而且我在过滤或策展其他社交媒体习惯方面非常谨慎。Facebook 是否仍然充斥着“虾耶稣”,还是那只是 2024 年的事情?我听说假视频中可爱动物被救助是最近的趋势。

That said... I don’t use Facebook, and I’m pretty careful at filtering or curating my other social media habits. Is Facebook still flooded with Shrimp Jesus or was that a 2024 thing? I heard fake videos of cute animals getting rescued is the latest trend.

垃圾内容问题很可能是一股不断增长的浪潮,而我却天真地没有意识到。

It’s quite possible the slop problem is a growing tidal wave that I’m innocently unaware of.

数据中心变得极不受欢迎的一年 The year that data centers got extremely unpopular

我差点跳过今年关于 AI 环境影响的写作(这是我在 2024 年写的),因为我不确定今年我们是否学到了什么_新_东西——AI 数据中心继续消耗大量能源,建设它们的军备竞赛也在以不可持续的方式加速。

I nearly skipped writing about the environmental impact of AI for this year’s post (here’s what I wrote in 2024) because I wasn’t sure if we had learned anything _new_ this year—AI data centers continue to burn vast amounts of energy and the arms race to build them continues to accelerate in a way that feels unsustainable.

2025 年有趣的是,公众舆论似乎正急剧转向反对新建数据中心。

What’s interesting in 2025 is that public opinion appears to be shifting quite dramatically against new data center construction.

以下是 12 月 8 日《卫报》的一个标题:超过 200 个环保组织要求暂停美国新建数据中心。地方层面的反对声似乎也在全面急剧上升。

Here’s a Guardian headline from December 8th: More than 200 environmental groups demand halt to new US datacenters. Opposition at the local level appears to be rising sharply across the board too.

我被 Andy Masley 说服,认为水资源使用问题大多被夸大了,这主要是一个问题,因为它分散了对能源消耗、碳排放和噪音污染等真正问题的注意力。

I’ve been convinced by Andy Masley that the water usage issue is mostly overblown, which is a problem mainly because it acts as a distraction from the very real issues around energy consumption, carbon emissions and noise pollution.

AI 实验室继续寻找新的效率提升方式,以更少的每 token 能耗提供更高质量的模型,但其影响是经典的杰文斯悖论——随着 token 变得更便宜,我们找到了更密集的使用方式,比如每月花费 200 美元使用数百万个 token 来运行编码智能体。

AI labs continue to find new efficiencies to help serve increased quality of models using less energy per token, but the impact of that is classic Jevons paradox—as tokens get cheaper we find more intense ways to use them, like spending $200/month on millions of tokens to run coding agents.

我的年度词汇 My own words of the year

作为一个痴迷于收集新词的人,以下是我在 2025 年最喜欢的词汇。你可以在我的定义标签中看到更长的列表。

As an obsessive collector of neologisms, here are my own favourites from 2025. You can see a longer list in my definitions tag.

* 氛围工程——我仍在犹豫是否应该尝试推广这个词!

* Vibe engineering—I’m still on the fence of if I should try to make this happen!

* 致命三重奏,这是我今年尝试创造的一个词,似乎已经扎根了。

* The lethal trifecta, my one attempted coinage of the year that seems to have taken root .

* 上下文腐烂,来自 Hacker News 上的 Workaccount2,指的是在会话过程中,随着上下文变长,模型输出质量下降的现象。

* Context rot, by Workaccount2 on Hacker News, for the thing where model output quality falls as the context grows longer during a session.

* 上下文工程,作为提示工程的一种替代,强调设计输入给模型的上下文的重要性。

* Context engineering as an alternative to prompt engineering that helps emphasize how important it is to design the context you feed to your model.

* 垃圾蹲守,由 Seth Larson 提出,指大语言模型幻觉出一个不正确的包名,然后被恶意注册以传播恶意软件。

* Slopsquatting by Seth Larson, where an LLM hallucinates an incorrect package name which is then maliciously registered to deliver malware.

* 氛围抓取——我的另一个没怎么流行起来的词,指由提示驱动的编码智能体实现的抓取项目。

* Vibe scraping—another of mine that didn’t really go anywhere, for scraping projects implemented by coding agents driven by prompts.

* 异步编码智能体,用于 Claude for Web / Codex 云 / Google Jules。

* Asynchronous coding agent for Claude for web / Codex cloud / Google Jules

* 提取性贡献,由 Nadia Eghbal 提出,指开源贡献中“审查和合并该贡献的边际成本大于对项目生产者的边际收益”的情况。

* Extractive contributions by Nadia Eghbal for open source contributions where “the marginal cost of reviewing and merging that contribution is greater than the marginal benefit to the project’s producers”.

2025 年总结 That’s a wrap for 2025

如果你读到了这里,希望你觉得这篇文章有用!

If you’ve made it this far, I hope you’ve found this useful!

你可以通过订阅源或电子邮件订阅我的博客,或者在 Bluesky、Mastodon 或 Twitter 上关注我。

You can subscribe to my blog in a feed reader or via email, or follow me on Bluesky or Mastodon or Twitter.

如果你希望每月收到这样的回顾,我还运营一份每月 10 美元的赞助者专属通讯,汇总过去 30 天 LLM 领域的关键进展。以下是 9 月、10 月和 11 月的预览版——我将在明天某个时候发送 12 月的版本。

If you’d like a review like this on a monthly basis instead I also operate a $10/month sponsors only newsletter with a round-up of the key developments in the LLM space over the past 30 days. Here are preview editions for September, October, and November—I’ll be sending December’s out some time tomorrow.

发布于 2025 年 12 月 31 日晚上 11:50 · 在 Mastodon、Bluesky、Twitter 上关注我,或订阅我的通讯

Posted 31st December 2025 at 11:50 pm · Follow me on Mastodon, Bluesky, Twitter or subscribe to my newsletter

近期文章 More recent articles

* Kimi K3,以及我们仍能从鹈鹕基准测试中学到什么 - 2026 年 7 月 16 日

* Kimi K3, and what we can still learn from the pelican benchmark - 16th July 2026

* 新的 GPT-5.6 系列:Luna、Terra、Sol - 2026 年 7 月 9 日

* The new GPT-5.6 family: Luna, Terra, Sol - 9th July 2026

* sqlite-utils 4.0,现已支持数据库模式迁移 - 2026 年 7 月 7 日

* sqlite-utils 4.0, now with database schema migrations - 7th July 2026

这是 2025 年:Simon Willison 发布的 LLM 年度总结,发布于 2025 年 12 月 31 日。

This is 2025: The year in LLMs by Simon Willison, posted on 31st December 2025.

1. AI 工程中的开放问题 - 2023 年 10 月 17 日,下午 2:18

1. Open questions for AI engineering - Oct. 17, 2023, 2:18 p.m.

2. 我们在 2023 年关于 AI 弄明白的事情 - 2023 年 12 月 31 日,晚上 11:59

2. Stuff we figured out about AI in 2023 - Dec. 31, 2023, 11:59 p.m.

3. 我们在 2024 年关于 LLM 学到的东西 - 2024 年 12 月 31 日,下午 6:07

3. Things we learned about LLMs in 2024 - Dec. 31, 2024, 6:07 p.m.

4. 过去六个月在 LLM 领域,用骑自行车的鹈鹕图解 - 2025 年 6 月 6 日,晚上 8:42

4. The last six months in LLMs, illustrated by pelicans on bicycles - June 6, 2025, 8:42 p.m.

5. 2025:LLM 之年 - 2025 年 12 月 31 日,晚上 11:50

5. 2025: The year in LLMs - Dec. 31, 2025, 11:50 p.m.

ai 2,129 openai 430 generative-ai 1,882 llms 1,849 anthropic 311 gemini 191 ai-agents 111 pelican-riding-a-bicycle 127 vibe-coding 92 coding-agents 224 ai-in-china 98 conformance-suites 11

ai 2,129openai 430generative-ai 1,882llms 1,849anthropic 311gemini 191ai-agents 111pelican-riding-a-bicycle 127vibe-coding 92coding-agents 224ai-in-china 98conformance-suites 11

上一篇:How Rob Pike 如何被一场 AI 垃圾“善意之举”骚扰

Previous:How Rob Pike got spammed with an AI slop "act of kindness"

月度简报 Monthly briefing

每月赞助我 10 美元,即可获得一份精心策划的邮件摘要,涵盖本月最重要的大语言模型发展动态。

Sponsor me for $10/month and get a curated email digest of the month's most important LLM developments.

互动版:图/公式 + 针对本篇提问 →