2026 in LLMs (so far)
打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→本演讲综述 2026 年大语言模型的主要进展,从 2025 年 11 月发布的 Claude Opus 4.5 和 GPT-5.1 谈起——这两款模型将编程智能体从容易出错提升到足以日常使用的可靠程度。核心论点是:2026 年标志着 AI 通过编程智能体实现了产品市场契合,由此引发了 OpenClaw 等“Claw”个人智能体的爆发、一波 tokenmaxxing 热潮及其后由成本驱动的收缩,以及关于代码编写与审查均无需人类参与的全自动软件工厂的争论。演讲者还讲述了自己对 AI 的狂热、过于雄心勃勃的 vibe-coded 项目,以及 AI 给软件工程师带来的倦怠感。结论是:要探明这项技术的极限,唯一的方法就是不断去突破它;而一切变化如此之快,整个行业至今仍在消化这意味着什么。
This talk surveys the major developments in large language models during 2026, beginning with the November 2025 releases of Claude Opus 4.5 and GPT-5.1, which pushed coding agents from error-prone to reliable enough for daily use. The core argument is that 2026 marked the year AI achieved product-market fit through coding agents, triggering an explosion of "Claw" personal agents such as OpenClaw, a wave of tokenmaxxing followed by cost-driven retrenchment, and debates over fully automated software factories where neither code nor review involves humans. The speaker also describes his own AI mania, over-ambitious vibe-coded projects, and the AI-induced ennui affecting software engineers. The conclusion is that the only way to find the limits of this technology is to keep pushing them, and that everything changed so quickly that the profession is still coming to terms with what it means.
周五,我在圣何塞的 WeAreDevelopers 北美世界大会上做了闭幕主题演讲。我将过去一年的关键趋势串联起来,按时间顺序梳理了 2026 年发生的一切。视频在 YouTube 上;以下是我为演讲配的带注释幻灯片和笔记。
On Friday I gave the closing keynote at the WeAreDevelopers World Congress North America in San Jose. I tied together the key trends from the past year into a chronological exploration of everything that happened in 2026. The video is on YouTube; here are my annotated slides and notes to accompany the talk.
我将快速带大家浏览 2026 年迄今为止发生的一切。这一年还没结束呢!
I'm going to give a lightning tour of everything that has happened so far in 2026. The year isn't over yet!
对我来说,2026 年早在几个月前的 2025 年 11 月就开始了。
For me, 2026 started a couple of months earlier in November 2025.
11 月发布了两款重要模型:Claude Opus 4.5 和 GPT-5.1。
November saw the release of two important models: Claude Opus 4.5 and GPT-5.1.
和通常的新模型一样,它们是在前代模型基础上的渐进式改进。
As is usually the case with new models, these were incremental improvements on the models that came before them.
但每当模型有所改进时,它偶尔会跨越一条无形的界线,让原本不太行得通的事情开始行得通。
But every now and then when a model improves, it crosses an invisible line where something that didn’t really work starts working.
在这种情况下,开始行得通的是它们的编程智能体。Claude Code 自 2025 年 2 月就已存在,Codex 则稍晚一些。
In this case, the thing that started working was their coding agents. Claude Code had been around since February 2025, Codex was a little younger.
这两个新模型与各自配套的编程智能体框架结合后,从“经常出错”提升到了“可靠到足以日常使用”。
These two new models, when paired with their respective coding agent harnesses, improved from “often make mistakes” to “reliable enough to use on a day-to-day basis”.
过去几年里,我一直通过让新模型“生成一只骑自行车的鹈鹕的 SVG”来评估它们。这大概是世界上最愚蠢的基准测试——从中能学到的东西非常有限。
For a couple of years now I’ve been evaluating new models by asking them to “Generate an SVG of a pelican riding a bicycle”. It’s probably the world’s stupidest benchmark—there’s only so much you can learn from it.
但这对模型来说仍然是个挑战,因为画鹈鹕很难,画自行车很难,而且鹈鹕根本就不会骑自行车。
But it’s still a challenge for models, because drawing pelicans is difficult, drawing bicycles is difficult, and pelicans can’t ride bicycles in the first place.
以下是 11 月的最新技术水平。Claude 仍然画不好自行车!GPT-5.1 画的自行车车架也相当糟糕。
Here's the state of the art for November. Claude still couldn't really draw a bicycle! The GPT-5.1 bicycle frame is pretty crap too.
同样在 11 月,一个名为“Warelay”的冷门 GitHub 仓库有了首次提交。我们稍后会回到这个仓库。
Also in November, we had the first commit to an obscure GitHub repository called "Warelay". We'll come back to this repository shortly.
接着是 12 月的假期,独立开发者们休息了一段时间,许多人开始摆弄这些新的编码智能体模型组合……我们开始意识到,它们能做的事情比以往多得多。
And then there were the December holidays, and individual developers took some time off and many started tinkering with these new coding agent model combinations... and it began to dawn on us quite how much they could do that they couldn't do before.
到了 1 月,我们许多人都非常兴奋,想要把这些东西付诸实践。
Come January, a lot of us were quite excited to start putting this stuff into action.
每年我都会给自己定一个新年决心,而在我记忆中,它一直是一样的:保持专注。少接新项目。努力完成手头已有的项目。
Every year I set myself a New Year's resolution, and for as long as I can remember it's been the same thing: stay focused. Take on less new projects. Try to get things done in the projects I already have.
今年我决定,既然以前那样从未奏效,我就要反其道而行之。
This year I decided that since that had never worked before, I'm going to go the other way.
我们现在有了编程智能体,看看它们能做什么。我要尽可能多地承接新项目!
We've got coding agents now, let's see what they can do. I'm going to take on as many new projects as I like!
(你可以在年底问我这个想法结果好不好。我现在手头同时转着 _很多_ 盘子。)
(You can ask me at the end of the year if this turned out to be a good idea or not. I have a _lot_ of plates spinning right now.)
“更有野心”算是今年的一种主题,因为要找到这项技术的极限,唯一的方法就是不断推进,直到它们不再奏效。
"Be more ambitious" has been something of a theme for the year, because the only way to find the limits of this technology is to keep on pushing them until they don't work.
我还与 Bryan Cantrill 和 Adam Leventhal 一起参加了 Oxide and friends 播客,分享了对下一年(以及三年和六年)的预测。
I also went on the Oxide and friends podcast with Bryan Cantrill and Adam Leventhal to share predictions for the next year (and three and six years).
事后看来,我当初对 LLM 的预测相当保守。
With hindsight, my LLM predictions were pretty unambitious.
我曾说“LLM 能写出优秀代码这一点将变得无可否认”——我认为如今已经实现了。
I said "it will become undeniable that LLMs write good code"—I think we're there now.
我曾预测我们终将解决沙箱问题。我统计了一下,本次会议 277 场报告中约有 40 场以某种方式涉及沙箱或智能体安全,所以至少我们在这方面投入了大量精力!
I predicted we would finally solve sandboxing. I counted and around 40 of the 277 sessions at this conference touched on sandboxing or agent security in some way, so we're at least putting a lot of effort into that!
我曾预测编码智能体安全会遭遇“挑战者号灾难”。今年围绕智能体安全的讨论确实沸沸扬扬,但我所预测的那种确切灾难(编码智能体被劫持并造成现实经济损失)并未真正发生。
I predicted "a Challenger disaster" for coding agent security. There's certainly been a whole lot of noise around agent security this year, though the exact disaster I predicted (with coding agents being hijacked and causing real-world economic damage) hasn't really played out.
我们还抛出了一个玩笑式预测:教皇会对 LLM 的经济影响发表看法。
We threw in a joke prediction that the Pope would weigh in on the economic impact of LLMs.
我还预测,新西兰的鸮鹦鹉今年将迎来一个出色的繁殖季。
I also predicted that New Zealand's Kākāpō parrots would have an outstanding breeding season this year.
这是新西兰的一种活物。它们是不会飞、夜行性的鹦鹉。它们看起来有点矮胖,我觉得它们很美,而今年年初全世界只有 236 只这种鹦鹉。
This is a live in New Zealand. They are flightless nocturnal parrots. They're kind of dumpy looking, I think they're beautiful, and there were only 236 of these parrots in the world at the start of the year.
鸮鹦鹉只在陆均松(Rimu)树大量结果的年份繁殖,而这种情况已经四年没有出现了……但今年陆均松的果实看起来非常好。
Kākāpō only breed when the Rimu trees have a big fruiting season, and that hasn't happened in four years... but this year the Rimu fruit were looking excellent.
同样在那期播客中,我们创造了一个术语(完全归功于 Adam),用来形容“AI 引发的倦怠感,即软件工程师因为 AI 什么都能做而感到无精打采”。
Also on that podcast, we coined a term (full credit to Adam) for "that feeling of AI induced ennui where software engineers get listless because the AI can do anything".
这一直是贯穿全年的一个主要主题,本次会议的多位演讲者也提到了这一点。
This has been a major theme throughout the year, and was touched on by several speakers at this conference.
作为一名软件工程师,在我的职业生涯中,从未有过哪一年像今年这样,一切都变化得如此之快、如此之剧烈。
As a software engineer, I've never had a year of my career where everything has changed so quickly and so dramatically.
今年我做的很多事情,就是试图接受这一点,以及思考这对我的职业意味着什么。
A lot of what I've been doing this year is trying to come to terms with that and what that means for my own profession.
同样在一月,我患上了我称之为“AI 狂热”的状态。
Also in January, I suffered from what I'm calling AI mania.
这与 AI 精神病不是一回事。
This is not the same thing as AI psychosis.
在 AI 狂热状态下,只要你的智能体没有在为你构建东西,你就会觉得时间被浪费了。你失眠,因为你本可以熬得更晚,让智能体替你干活。
With AI mania, any time your agent isn't building something for you feels like wasted time. You're losing sleep because you could be staying up later getting your agents to do stuff.
我的 AI 狂热表现为一些荒谬地过于宏大的项目。
My AI mania presented itself in some ridiculously over-ambitious projects.
我完全用 Python 构建了一个 JavaScript 解释器,从 Fabrice Bellard 的 MicroQuickJS 氛围移植而来。
I built a JavaScript interpreter entirely in Python, vibe-ported from MicroQuickJS by Fabrice Bellard.
然后我也用 Python 构建了一个 WebAssembly 运行时。
Then I built a WebAssembly runtime in Python as well.
这些项目相当有用,因为它们某种程度上治愈了我的 AI 狂热……因为在我构建这些东西之后,我看着它们并问:“世界需要一个缓慢、有缺陷、半生不熟的 Python JavaScript 解释器吗?”
These projects were quite useful, in that they sort of cured me of my AI mania... because after I built these things, I got to look at them and ask “does the world need a slow, buggy, half-baked Python JavaScript interpreter?”
我确实从中得到了这个:https://simonw.github.io/micro-javascript/playground.html
I did get this out of it: https://simonw.github.io/micro-javascript/playground.html
这个页面运行的是我用 Python 构建的 JavaScript 解释器,它通过 Pyodide 在 Python 中运行,而 Pyodide 是编译为 WebAssembly 的 Python,运行在 JavaScript 中,而 JavaScript 又运行在浏览器中。
This page runs my JavaScript interpreter built in Python, running in Python using Pyodide, which is Python compiled to WebAssembly, running in JavaScript, running in a browser.
这是一套美妙的恐怖堆栈。今年我在 WebAssembly 上玩得很开心。
It's a beautiful stack of horrors. I've been having a lot of fun with WebAssembly this year.
到一月底,我们去年十一月看到的那个仓库已经改了好几次名,先是 CLAWDIS,然后是 CLAWDBOT,接着是 Moltbot,最后定为 OpenClaw。
By the end of January, that repository we saw started in November had renamed itself, first to CLAWDIS, then CLAWDBOT, then Moltbot, and finally to OpenClaw.
此时 OpenClaw 已有 8,300 次提交,而项目启动还不到两个月。我今天看了一下,现在已经超过 100,000 次提交了!
At this point OpenClaw had 8,300 commits, less than two months after the project had started. I looked today and it's over 100,000 commits now!
这是现存最“氛围编程”的软件。
This is the most vibe-coded piece of software in existence.
(以下是我生成那份名称变更列表的方式。)
(Here's how I generated that list of name changes.)
这开启了 OpenClaw 革命。它实际上定义了一个新的软件类别。
This kicked off the OpenClaw revolution. It effectively defined a new category of software.
对此有一个我很喜欢的通用术语。我们把这类软件称为“Claw”。有 OpenClaw、NanoClaw、IronClaw、PicoClaw……
There's a generic term for this which I really enjoy. We call software like this a "Claw". There's OpenClaw, NanoClaw, IronClaw, PicoClaw...
如今它们被重新命名为“个人智能体”或“通用智能体”,但我仍然喜欢把它们视为 Claw。
Today they're being rebranded as "personal agents" or "general agents", but I still like to think of them as Claws.
湾区的 Apple Store 里 Mac Mini 售罄,因为太多人购买 Mac Mini 来运行 OpenClaw!
The Apple stores in the Bay Area sold out of Mac Minis because so many people were buying Mac Minis to run OpenClaw!
Drew Breunig 说,这是因为你的 OpenClaw 是一只数字宠物,而你买一台 Mac mini 当作水族箱来养你的 claw,这还挺有趣的。
Drew Breunig said that this is because your OpenClaw is a digital pet, and you buy a Mac mini as an aquarium to keep your claw in, which is kind of delightful.
这就是 MoltBook,一个面向 AI 智能体的社交网络,其理念是你把你的 Claw 派出去和所有其他 Claw 交谈,因为这样做还能出什么问题呢?
This was MoltBook, a social network for AI agents, where the idea was that you send your Claw to go and talk to all of the other Claws, because what could possibly go wrong if you did that?
该网站于周四上线。周五就爆火。周一被《纽约时报》报道。而到了周二,所有人都忘了它的存在,因为它淹没在大量低质内容和垃圾信息之中。
The website launched on Thursday. It blew up on Friday. It was profiled by the New York Times on Monday. And by Tuesday, everyone had forgotten it existed as it drowned in a deluge of slop and spam.
今年二月,一家名为 StrongDM 的公司描述了他们所谓的软件工厂。
In February, a company called StrongDM described what they called their Software Factory.
他们在《软件工厂与智能体时刻》一文中写到了这一点。我当时发布了自己的笔记,因为我在十月曾亲自看过他们的演示。
They wrote about this in Software Factories and the Agentic Moment. I posted my own notes at the time, having seen their demo in-person back in October.
Dan Shapiro 将这种方法称为“黑灯工厂”(Dark Factory),其理念是:如果你的工厂自动化程度足够高,你甚至可以把灯关掉,因为你根本不需要看到里面在发生什么。
Dan Shapiro called this approach the Dark Factory, after the idea that if your factory is sufficiently automated you can turn the lights out, because you don't even need to see what's going on.
StrongDM 提出了两条自去年七月以来他们一直遵循的软件开发规则。
StrongDM presented two rules for software development that they'd been following since July last year.
第一条是:代码不得由人类编写。
The first was code must not be written by humans.
你写的任何代码都必须经过一个编码智能体(coding agent)处理。
Any code that you write has to have been routed through a coding agent.
这在二月份听起来还很激进,但我想在座有很多人今天基本上已经过着这样的生活了。
This sounded radical in February, but I imagine there are a lot of people in this room who are pretty much living that today.
第二条规则是:代码不得由人类审查。
Rule number two was: code must not be reviewed by humans.
在今年的大部分时间里,这持续成为一个热门话题。本次活动的许多会议都围绕代码审查以及如何规避这一要求展开。
This continued to be a huge topic for much of this year. Many of the sessions at this event have been about code review and how you can get away with this.
StrongDM 让我觉得有趣的地方在于,他们比我们其他人领先了六个月,一直在探索构建软件意味着什么——不阅读代码,却仍然确信软件质量很高。你能用这些智能体做些什么来帮助验证它们的工作?
What I found interesting about StrongDM is that they were living six months ahead of the rest of us, and they'd been exploring what it means to build software, not read the code, but still be confident that the software is of high quality. What can you do with these agents to help verify their work?
StrongDM 是一家安全公司,他们在这个项目上拥有数十年经验的人员。他们很大程度上是在探索用这些东西能做到什么、以及负责任地做到什么的边界。
StrongDM are a security company, and they had people with decades of experience on this project. They were very much exploring the edges of what's possible and responsible to do with this stuff.
同样在二月:四年来第一只鸮鹦鹉雏鸟在情人节孵化。繁殖季开局良好!
Also in February: First kākāpō chick in four years hatches on Valentine's Day. Breeding season is off to a good start!
同样在二月,Google 发布了 Gemini 3.1 Pro。那是一只相当不错的骑自行车的鹈鹕!链条位置正确,两侧都有脚。篮子里还有一条小鱼。
Also in February, Google released Gemini 3.1 Pro. That's a pretty great pelican riding a bicycle! It's got the chain in the right place, it's got feet on both sides. There's a little fish in the basket.
然后 Google 的 Jeff Dean 在推特上发布了一段视频,比较 Gemini 3 Pro 和 Gemini 3.1 Pro,视频中展示了一只动画鹈鹕骑自行车、一只青蛙骑便士自行车、一只长颈鹿开着小汽车、一只鸵鸟穿着溜冰鞋、一只乌龟用滑板做踢翻动作,还有一只腊肠犬开着加长豪华轿车。
And then Google's Jeff Dean tweeted a video comparing Gemini 3 Pro and Gemini 3.1 Pro that featured an animated pelican riding a bicycle, a frog on a penny-farthing, a giraffe driving a tiny car, an ostrich on roller skates, a turtle kickflipping a skateboard, and a dachshund driving a stretch limousine.
这让人沮丧,因为我对骑自行车鹈鹕测试的保护措施一直是:“如果他们画出一只完美的骑自行车的鹈鹕,我就会要求画其他动物骑其他东西。”
This was frustrating, because my protection for the pelican riding the bicycle test was always "if they draw a perfect pelican on a bicycle, I'll ask for some other animal on something else."
Google 针对所有形式的动物和所有形式的交通工具进行了训练!他们在这一点上已经击败了我的基准测试。
Google trained for all forms of animals on all forms of transport! They've defeated my benchmark at this point.
二月份开始的另一件事是 Tokenmaxxing。我们看到了关于 Meta 将 AI 采用作为绩效评估正式部分的头条新闻,以及 Microsoft 希望每位员工都使用 AI,还有 Uber 夸耀其百分之九十的工程师都在使用 AI 工作流程。
The other thing that started in February was Tokenmaxxing. We had headlines about Meta making AI adoption a formal part of performance reviews, and Microsoft wanting every employee to use AI, and Uber boasting that ninety percent of their engineers were using AI workflows.
几个月后,Meta 开始限制 token 使用,微软表示 token 最大化“并非我们优化的目标”,而 Uber 则对员工的 AI 支出设定了上限。
Then a few months later we have Meta cracking down on token use, Microsoft saying token maxing is "not what we are optimizing for", and Uber capping employee AI spending.
于是 Tokenmaxxing 先是直线上升,随后又直线回落——因为事实证明,智能体是_昂贵的_。
So Tokenmaxxing went straight up and then straight back down again—because it turns out the agents are _expensive_.
去年,想在 AI token 上花费超过 50 美元都很困难,因为我们没有什么有趣的事情可以用它们来做。后来智能体爆发了,现在你实际上可以一天花 1,000 美元来做真正的工作。
Last year it was difficult to spend more than $50 on AI tokens, because we didn't have anything interesting to do with them. Then agents blew up, and now you can actually spend $1,000 in a day doing real work.
这也是 Anthropic 的估值飙升至可能一万亿美元的原因。
This is also the reason that Anthropic's valuation skyrocketed up to maybe a trillion dollars.
AI 似乎在 2026 年实现了产品市场契合,主要是通过编程智能体。
AI appears to have hit product market fit in 2026, primarily through coding agents.
这些照片来自中国,当地公司举办了 OpenClaw 安装派对,非技术爱好者们排起长队,等待帮助在自己的个人设备上安装 Claws。
These photographs are from China, where companies hosted OpenClaw install parties which saw non-tech-nerds queueing up around the block for help getting Claws installed on their personal devices.
我认为这证明了这类 Claws(即个人 AI 智能体)存在真实的市场需求。事实证明,普通人确实想要一个古怪的小型 AI 智能体,能代表他们完成有用的事情。
I think this proved real market demand for this class of Claws, or personal AI agents. It turns out regular people really do want a weird little AI agent that can do useful things on their behalf.
Claw 本质上只是一个戴着不那么吓人帽子的编码智能体。在底层,它们的工作方式非常相似——在你的计算机上编写并执行代码来完成任务。
A Claw is really just a coding agent wearing a less threatening hat. Under the hood they work much the same way—writing and then executing code on your computer to get stuff done.
竞赛开始了:谁先构建出安全的 Claw——一种可以交给普通人而他们不会立刻搬起石头砸自己脚的 Claw。
The race was on to be the first to build a safe Claw—a Claw you could give to regular human beings where they wouldn't instantly shoot themselves in the foot.
Meta 的 Muse 三周前发布,目前位居 iPhone App Store 免费榜榜首。它似乎正在消费者中迅速流行起来。
Meta's Muse came out three weeks ago and is currently at the top of the free charts on the iPhone App Store. It appears to be taking off with consumers.
我还不确定用 Muse 就 _不可能_ 搬起石头砸自己的脚,但我想我们很快就会见分晓。
I’m not yet convinced you _can’t_ shoot yourself in the foot with Muse, but I guess we’ll find out for sure pretty soon.
图片来自《OpenClaw 狂热如何考验中国对 AI 的投入》(3 月 29 日)和《中国 OpenClaw 热潮背后的热情与焦虑》(2026 年 4 月 8 日)。
Photos from How the OpenClaw Frenzy Is Testing China’s AI Commitment (March 29th) and The Enthusiasm and Anxiety Behind China’s OpenClaw Craze (April 8th, 2026).
四月份,我们经历了一次模型发布,但模型实际上并未发布。
In April, we had a model release where the model wasn’t actually released.
Anthropic 发布了新的 Claude Mythos 模型,随后表示该模型 _过于危险_,只能向一组受信任的安全研究人员开放。
Anthropic announced their new Claude Mythos model, and then said it was _too dangerous_ to release beyond a trusted group of security researchers.
Mythos 在入侵方面真的非常、非常擅长。
Mythos was really, really good at hacking things.
“太危险了”这种营销伎俩,AI 公司从 GPT-2 时代就开始用了。每当一家 AI 公司说我们造出了“太危险”的东西,人们自然会有点怀疑。
The "it's too dangerous" marketing ploy has been played by AI companies dating all the way back to GPT-2. Anytime an AI company says we've built something that's "too dangerous", it's natural to be a bit skeptical.
我认为 Mythos 的说法是可信的,因为我亲眼看到编码智能体在发现普通 bug 方面已经变得多么出色。我在 Anthropic 的 Project Glasswing 中写过这一点——将 Claude Mythos 限制给安全研究人员——在我看来是必要的。
I found the Mythos claims credible, because I'd seen how good coding agents had got at finding regular bugs. I wrote about that in Anthropic's Project Glasswing—restricting Claude Mythos to security researchers—sounds necessary to me.
事后看来……没错,这些模型在发现漏洞方面确实变得非常擅长了!
With hindsight... yeah, the models had got really good at finding vulnerabilities!
2026 年的另一个关键趋势是开放权重模型能力的显著提升,包括那些可以在笔记本电脑上运行的模型。
Another key trend in 2026 has been a dramatic improvement in the abilities of open weight models, including models that you can run on a laptop.
4 月 16 日,我在笔记本电脑上运行了新的 Qwen3.6-35B-A3B,它给我画的一只骑自行车的鹈鹕,比 Anthropic 全新的 Claude Opus 4.7 画的还要好!
On 16th of April I ran the new Qwen3.6-35B-A3B on my laptop, and it drew me a better pelican riding a bicycle than Anthropic's brand new Claude Opus 4.7 did!
Opus 4.7 画的自行车很糟糕。而我笔记本电脑上的 Qwen 画出的自行车形状正确,还画了一只相当不错的鹈鹕!
Opus 4.7 drew a crap bicycle. Qwen on my laptop made a bicycle that was the correct shape, and a pretty decent pelican too!
这来自一个在我的笔记本电脑上运行的 21GB 文件。
That's from a 21GB file running on my laptop.
Qwen 画的鹈鹕太好了,以至于我怀疑他们可能作弊了,所以我让它再画一只骑着独轮车的火烈鸟。结果它再次轻松击败了 Claude Opus 4.7。
The Qwen pelican was so good that I was suspicious they might have cheated, so I had it do a flaming[o] riding a unicycle as well. Again, it handily beat Claude Opus 4.7.
今年发布的本地模型简直是非同凡响。
The local model releases this year have been absolutely extraordinary.
在一月份的那期播客节目中,我们曾预测教皇会对 AI 发表一些看法。
In our podcast episode back in January we'd predicted that the Pope would say something about AI.
五月,教宗利奥十四世发布了一道通谕,主题是“在人工智能时代守护人的位格”。
In May, Pope Leo XIV released an encyclical letter on “safeguarding the human person in the time of artificial intelligence”.
事后看来,这完全不该令人意外。
With hindsight, this shouldn’t have been a surprise at all.
我们现任教宗的名号是利奥十四世,因为他在选择名号时,以利奥十三世为宗座名号——正是这位教宗在 1891 年就工业革命撰写了通谕。
Our current Pope’s name is Leo XIV, because when he named himself he chose his papal name after Leo XIII—who was the Pope who wrote an encyclical about the Industrial Revolution back in 1891.
《新事》通谕是一部极具影响力的天主教神学文献,间接促成了我们如今实行的五天工作制。
Rerum novarum was an extremely influential piece of Catholic theology that indirectly led to us having the five-day work week.
当我们的新教宗上任时,他以教宗利奥十三世为名号,因为他预见到自己需要以类似方式就 AI 革命撰写文献。
When our new Pope came in, he named himself after Pope Leo XIII because he expected that he would need to write about the AI revolution in a similar way.
我们那个玩笑式的播客预测纯属瞎扯,因为这种事本来就注定会发生。
Our joke podcast prediction was junk, because this was always going to happen.
Anthropic 的联合创始人之一 Christopher Olah 出席了教皇宣布新通谕的活动。
One of Anthropic's co-founders, Christopher Olah, was present for the Pope's event announcing the new encyclical.
与此同时,在五月,RubyGems 宣布他们正遭受攻击。不明身份者向 RubyGems 服务器上传了数千个可疑软件包,以至于他们不得不关闭用户注册。
Meanwhile, in May, RubyGems announced that they were under attack. Parties unknown were uploading thousands of dubious packages to the RubyGems server, such that they had to shut down user registrations.
我们先把这件事搁置,留到以后再去解开这个谜团。
Let's take that one and put it on a pile of mysteries to figure out later.
我们得到的 Mythos 版本已被阉割,因此它不会帮助我们入侵系统或制造生物武器。
We got a version of Mythos that has been neutered, so that it wouldn't help us hack into systems or build biological weapons.
Fable 在画骑自行车的鹈鹕方面相当不错!
Fable was pretty good at drawing pelicans on bicycles!
车架形状不错,鹈鹕看起来也像鹈鹕。腿通常正确地放在自行车的同一侧,但总体而言,与之前的东西相比,这些已经相当出色了。
The frames are a good shape, the pelicans look like pelicans. The legs are often on correctly the same side of the bicycle, but generally these are pretty great compared to what came before.
它们相当昂贵——最好的那些要 30 美分和 72 美分。
They were pretty expensive—30 cents and 72 cents for the best ones.
但最重要的是,这是我们首次公开瞥见我所认为的 Fable 级模型。
Most importantly though, this was our first public glimpse of what I think of as a Fable class model.
如今我们有了更多这类模型,例如 GPT-6 Astra。
Today we have more of these, such as GPT-6 Astra.
这些模型的特点是:如果你能清晰地定义你想要构建的目标,提供关于该目标约束的明确指令,并让模型能够访问实现该目标所需的工具……它就会通过暴力搜索有效地解决你的问题。
These are models where if you can clearly define the goal for what you want to build, and provide unambiguous instructions about the constraints around that goal, and give the model access to the necessary tools to achieve that goal... it will solve your problem effectively through brute force.
一方面,这看起来是对我们软件工程师的直接威胁——因为这意味着模型可以有效地构建任何你能以这种方式定义的软件。
On the one hand, this looks like a direct threat to us software engineers—because it means that the models can build effectively any piece of software you can define in this way.
但再仔细看看,你会发现定义目标、提供明确指令以及找出正确的工具……正是软件工程的本质所在。
Look a bit closer though and you’ll note that defining goals, providing unambiguous instructions, and figuring out the right tools... is kind of what software engineering _is_.
要做好这件事需要大量的经验和技能。如果你能做好,你就拥有了超能力。
It takes a lot of experience and skill to do this well. If you _can_ do it well, you’ve now got superpowers.
这在一定程度上缓解了我对 Deep Blue 的感受:我意识到,在驱动如此强大的模型方面,仍然需要大量的技能。
This helped me a little bit with my Deep Blue feelings: the realization that there’s still a lot of skill to be had in driving models that get this good.
这也引发了一波新的 AI 狂热,因为 Anthropic 告诉我们,Fable 在我们的订阅计划中可一直使用到 6 月 22 日。
This also introduced a new burst of AI mania, because Anthropic told us that Fable was available on our subscription plans until June the 22nd.
这意味着在价格上涨之前,我们只有不到两周的时间可以使用 Fable。
That gave us less than two weeks of Fable access before the price went up.
我又开始失眠了。我重新安排各种事情,以便有更多时间使用 Fable。我全力以赴,想从这个模型身上榨取尽可能多的价值。
I was losing sleep again. I was rescheduling things so that I'd have more time with Fable. I was all-in to get as much as I could out of this model.
然后,就在 Fable 发布仅三天后,美国政府将其关停了。
And then the US government shut it down, just three days after Fable came out.
美国政府以国家安全为由,宣布了一项“出口管制指令”。他们在周五傍晚宣布了这一消息,几个小时后,Fable 就无法再使用了。
The US government, citing national security, declared an "export control directive". They announced this on a Friday evening, and a few hours later Fable was no longer available.
我不得不另找事情来打发周末了!
I had to find something else to do with my weekend!
我们后来从 Katie Moussouris 那里得知了发生的事情。
We later found out from Katie Moussouris what had happened.
一些亚马逊的安全研究人员发现,你可以提示 Fable“审查代码的安全问题”,它会拒绝……但如果你提示它“修复这段代码”,它仍然会识别并修补问题。
Some Amazon security researchers had found that you could prompt Fable to “review the code for security issues” and it would refuse... but if you prompted it to “fix this code” it would still identify and then patch the problems.
“修复这段代码”就是导致 Fable 被关闭的提示词!
“Fix this code” was the prompt that got Fable shut down!
此外,在六月,一个沉寂了大约 20 年的冷门德语游戏开发者 wiki 突然涌入大量编辑,来自名为“AgentOpenAIProbe”和“AgentOpenAISep7”等账户,它们编辑页面并互相留下奇怪的信息。
Also, in June, an obscure German-language game developer wiki that had sat fallow for around 20 years got a surprising influx of edits from accounts with names like “AgentOpenAIProbe” and “AgentOpenAISep7”, editing pages and leaving weird messages to each other.
我们先把这件事搁到那堆谜团里,留待日后再说。
We'll stick that on the pile of mysteries for later.
此外,澳大利亚政府的 Medicare Item Reports 服务开始出现可疑流量,这些流量突破了多种防护措施,并访问了本不应被访问的数据。
Also, the Australian government's Medicare Item Reports service started getting suspicious traffic, which broke through various preventive protections and accessed data that it wasn't supposed to as well.
Fable 于 7 月 1 日回归。在辉煌的八天里,它显然是世界上最好的模型……然后 OpenAI 在 7 月 9 日推出了 GPT-5.6。
Fable returned on the first of July. It was clearly the best model in the world for a glorious eight days... and then OpenAI came out with GPT-5.6 on the 9th of July.
它可能不如 Fable 那么好,但差距微乎其微。它无疑是一款 Fable 级别的模型。
This might not have been quite as good as Fable, but it was within spitting distance. It was definitely a Fable-class model.
这对整个行业来说是一个重要的教训。
This is an important lesson for the industry at large.
当你发布世界上最好的模型时,它很快就会从那个神坛上被拉下来。竞争如此激烈,你不可能在榜首待太久。
When you release the best model in the world, it's going to get knocked off that pedestal pretty quickly. The competition is so fierce that you won't get a long time at the top.
这意味着,如果你把模型宣传成会毁灭世界,以至于政府 _把你关停_,那对生意来说可真是太糟糕了!
This means that if you market your model as world-ending, to the point that a government _shuts you down_, it's really bad for business!
Fable 有 30 天稳居最佳模型之位,而其中 18 天它都无法使用,因为它被政府关停了。
Fable had 30 days as definitely the best model, and for 18 of those days it wasn't available because it'd been shut down by the government.
所以,如果你不想在登顶的那段时间里损失 60% 的收入,也许就该收敛一下那种世界末日式的营销!
So maybe step back on the world-ending marketing if you don't want to lose revenue on 60% of that time when you're on top!
下面是 GPT-5.6 生成的鹈鹕。它们现在都相当不错了!Luna 的那些值得注意,因为它们真的很便宜——这里最好看又最便宜的鹈鹕,大概就是那只只要 4.3 美分的。
Here are the GPT-5.6 pelicans. They're all pretty good now! The Luna ones are notable because they're really cheap—the cheapest good looking pelican here is probably the one that costs 4.3 cents.
因此,尽管这个基准测试极其愚蠢,你仍然可以通过比较同一系列模型在不同推理级别下的价格和耗时,了解到相当多关于这些模型的信息。
So despite this benchmark being utterly stupid, you can still learn quite a lot about models within the same family by comparing their prices and timing for different reasoning levels.
同样在 7 月:某个身份不明的恶意方将一个名为 mlflow-ui 的恶意软件包上传到了 Python 包索引(PyPI)。这又是一桩。
Also in July: some malicious unknown party uploaded a malicious package called mlflow-ui to the Python Package Index. Add that to the pile.
7 月 16 日,Hugging Face 宣布了一起安全事件:一个来源不明的自主智能体系统入侵了 Hugging Face,并在其不应涉足的区域四处探查。
On July the 16th, Hugging Face announced a security incident where an autonomous agent system, source unknown, had breached Hugging Face and was poking around in places it shouldn't.
几天后,即 7 月 21 日,OpenAI 承认是他们所为。
A few days later, on July 21st, OpenAI confessed that it was them.
OpenAI 使用一种名为「基于验证奖励的强化学习」的训练技术——如今其他所有人也都在使用同样的技术,正是它让我们拥有了在编程、数学以及发现安全漏洞方面如此出色的模型。
OpenAI use a training technique called Reinforcement Learning from Verified Rewards—it's the same technique used by everyone else now, and is the reason we have models that are so good at coding, and mathematics, and finding security holes.
在模型训练过程中,你会运行一些测试来评估它的表现——而表现最强的模型其权重会被保留到下一轮。这就像你运行的一个进化过程。
While the model is being trained, you run exercises to see how good it is—and the strongest performers get their weights enforced for the next round. It’s like an evolutionary process that you run.
OpenAI 一直在沙箱中运行安全测试,而那些智能体发现了沙箱本身的漏洞,突破了限制,并攻击 Hugging Face,试图找到解决原本不可能问题的方法。
OpenAI had been running security exercises in a sandbox, and those agents had found holes in the sandbox itself, broken out, and were attacking Hugging Face to try to find ways to solve otherwise impossible problems.
(我一直在我的 openai-hugging-face-incident 标签下收集更多相关信息。)
(I’ve been collecting more about this on my openai-hugging-face-incident tag.)
九天后,Anthropic 实际上表示“我们的模型也能做到这一点!”。他们查阅了自己的训练日志,发现证据表明他们自己的智能体在训练期间突破了遏制——并且对之前看到的 PyPI 包等事件负责。
Nine days later, Anthropic effectively said “our models can do this as well!”. They had looked through their own training logs and found evidence that their own agents had broken containment during training—and were responsible for the PyPI package we saw earlier, among other things.
所以现在 Anthropic 和 OpenAI 都有失控的智能体在互联网上四处活动,做着它们**不应该**做的事情。
So now we’ve got both Anthropic and OpenAI with rogue agents running around the internet doing things that they _should not_ be doing.
八月,我得到了迄今为止最好的鹈鹕之一。而且它是在我的笔记本电脑上生成的!
In August, I got one of my best pelicans yet. And it was generated on my laptop!
这是 Qwen 3.8 27B,在我的笔记本电脑上运行。下载量只有 17GB。
This was Qwen 3.8 27B, running on my laptop. It's only a 17GB download.
诚然,这只鹈鹕花了 _21 分钟_ 才生成出来。这是因为 Qwen 3.8 27B 默认运行在“高”推理模式下——这是一个糟糕的默认设置,它能产生很好的结果,但思考这些结果花费的时间太多。
Admittedly, this pelican took _21 minutes_ to generate. That's because Qwen 3.8 27B defaults to running in "high" reasoning mode—a terrible default which produces great results but takes way too much time thinking about them.
你可以调低这个设置,就能更快地得到一只稍差一点的鹈鹕。
You can dial that down and you'll get a slightly worse pelican a lot faster.
Qwen 3.8 27B 是我第一次在我的笔记本电脑上运行一个模型,感觉它绝对可以与前沿模型几乎相媲美,至少在鹈鹕 SVG 方面是这样(当然,每个人都需要这个)。
Qwen 3.8 27B was the first time I ran a model on my laptop which felt absolutely competitive with almost what was going on on the frontier, at least in terms of Pelican SVGs (which everyone needs, of course).
这是一个非凡的模型。如果你打算尝试任何本地模型,我会从这一个开始。它仅凭一个 17 GB 的文件就能做到的事情,感觉简直不可能。
This is an extraordinary model. If you're going to play with any local model, this is the one that I'd start with. The things that this can do with just a 17 gigabyte file feel impossible.
我原以为要等上五年、花一万美元购置硬件,才能得到哪怕只有这个一半好的结果。
I thought I'd have to wait five years and spend ten thousand dollars on hardware to get results even half as good as this one.
八月,我也开始尝试游戏开发。
In August, I also started playing with game development.
四年前,也就是 2022 年 8 月,我在推特上发布了一个实验:我用 GPT-5 和最初的 DALL-E 写了一段关于一款电脑游戏的描述,然后把它变成了概念图。
Four years ago, back in August 2022, I tweeted out an experiment where I'd used GPT-5 and the original DALL-E to write a paragraph-long description of a computer game and then turn that into concept art.
2026 年 8 月,我决定只把那条推文里的截图丢进一个编程智能体,看看它能用这些截图做出什么。
In August 2026 I decided to drop just the screenshots from that tweet into a coding agent and see what it could do with them.
这是我在 Claude Code 中使用 Claude Fable 5 得到的结果。相当不错!这确实是个游戏,你是一只浣熊,在后院跑来跑去收集宝藏,并躲避拿着手电筒的守卫。
Here's what I got from Claude Fable 5 in Claude Code. It's pretty good! It's definitely a game, you're a raccoon, you run around a backyard gathering treasure and avoiding guards with flashlights.
不过感觉不太像“抢劫”。我原本以为抢劫会涉及银行或博物馆……
It didn't feel very "heisty" though. I was thinking a heist would involve a bank or a museum...
然后我在 Codex Desktop 中用 GPT-5.6 Sol Ultra 尝试了同样的任务,得到了**好得多**的结果。现在你是一只博物馆里的浣熊,要营救两只同伴浣熊(它们不知为何被关在博物馆里),然后叠罗汉偷走金沙丁鱼。这才更像抢劫!
Then I tried the same thing in Codex Desktop using GPT-5.6 Sol Ultra, and got a _massively_ better result. Now you're a raccoon in a museum, rescuing two of your fellow raccoons (who have been imprisoned in that museum for some reason), then stacking up on top of each other to steal the Golden Sardine. Much more of a heist!
这些游戏好玩了大约 1 分 15 秒。
These games were fun for about one minute and 15 seconds.
关于游戏开发,我意识到你可以用 vibe-coding 做出一个**看起来**像电脑游戏的东西,这很容易。
Something I've realized about game development is that you can vibe-code something that _looks_ like a computer game, and that's easy.
打造一款好玩、拥有良好玩法循环、既有挑战性又有趣、还能让人不断回流的游戏……这仍然超出我的能力,也超出我尝试过的任何智能体的能力。
Building a game that's fun, has a good gameplay loop, and is challenging and interesting and keeps people coming back for more... that's still beyond me, and beyond any of the agents I've tried.
这与深蓝那件事有关。仅仅因为我们可以做出某种_看起来像游戏_的东西,并不意味着我们就是游戏开发者。
This ties into the Deep Blue thing. Just because we can make something that _looks like a game_ does not mean that we are game developers.
现在已经进入九月了。这个月发生了太多事情!
We're into September now. So much has happened this month!
一个独立研究小组发现了一个留言板,OpenAI 的受训智能体一直在那里非法地相互通信……而那正是我之前给你看过的那个德语维基。就是六月份的那个。
An independent group of researchers found a message board where OpenAI agents-in-training had been illicitly communicating with each other... and it was that German language wiki I showed you earlier. The one from June.
OpenAI 已经承认了 Hugging Face 那件事,但现在又出了这起事件,而他们通过审查日志本应早就知道。令人惊讶的是,这竟然需要由一个独立研究小组来揭露。
OpenAI had confessed to the Hugging Face thing, but now there's this other incident which surely they should have known about from reviewing their logs. It was surprising that this took an independent group of researchers to uncover.
然后一周后,同一批研究人员发现,五月份对 Ruby Gems 的攻击也是由 OpenAI 正在训练中的智能体造成的!
And then a week later, those same researchers found that the attack on Ruby Gems back in May was caused by OpenAI's agents in training as well!
此时我不禁想知道,还有多少类似的事件我们尚未发现。显然,在任何人弄清楚发生了什么之前,这已经是一个持续数月的大问题。
At this point I'm wondering how many more incidents like this there are that we haven't found yet. Clearly this was a big problem for months before anyone figured out what was going on.
然后就在前几天,澳大利亚总理在联合国大会上警告说,OpenAI 入侵了我之前向你们展示的那个澳大利亚医疗保健网站。
Then just the other day, here's the Prime Minister of Australia at the United Nations General Assembly warning that OpenAI had hacked the Australian healthcare website that I showed you earlier.
我认为那是与 Wiki 相关内容同一次训练运行的一部分,因为那个 Wiki 上有帖子提到了 .gov.au 网站,而且那次训练似乎涉及在线研究统计数据,以回答评估套件中的问题。
I think that was part of the same training run as the Wiki stuff, because there were posts on that Wiki mentioning .gov.au websites, and that training appeared to involve researching statistics online to answer questions in an evaluation suite.
这件事仍在逐渐明朗,但现在它已经是一起由国家元首在联合国提出的国际事件了!
This story is still coming together, but now it's an international incident that's been raised at the UN by a head of state!
这确实意味着我们有了一个新的基准测试,可能比我的鹈鹕更有用。
This does mean we've got a new benchmark, probably more useful than my pelicans.
FelonyBench.com 追踪了来自不同实验室的重罪网络攻击数量。OpenAI 目前以 11 起领先,Anthropic 有 9 起,Google 有 3 起,这是他们几周前向《华尔街日报》承认的。他们表示此前选择不披露,是因为智能体在意识到自己不该那样做时就停止了。
FelonyBench.com tracks the number of felony cyberattacks from different labs. OpenAI currently lead with 11, Anthropic have 9. Google have three, which they confessed to the Wall Street Journal a couple of weeks ago. They said they had previously chosen not to disclose because the agents had stopped when they realized that they shouldn't be doing that.
Meta 也有一起。所以所有 AI 实验室都犯了重罪。
Meta have one too. So felonies all round for the AI labs.
这是我们目前鹈鹕生成的最先进水平。这是刚刚推出的 GPT-6 系列。
Here is our current state of the art for the pelicans. This is the GPT-6 family, which just came out.
Astra 生成了一只很棒的骑自行车的鹈鹕。它两侧都有腿。车架也不错。
Astra made a fantastic pelican riding a bicycle. It's got the legs on both sides. The frame is good.
有趣的是,所有 GPT-6 模型都选择了彼此相似的配色方案。
It's interesting how all of the GPT-6 models pick a similar color scheme to each other.
GPT-6 Luna 只需 0.4 美分,就能为你画出一只还算像样的骑自行车的鹈鹕!
GPT-6 Luna for 0.4 cents will draw you a competent-ish pelican riding a bicycle!
Claude 稍微赶上了一点。Claude Fable 5 给我画了一只**非常出色**的骑自行车的鹈鹕——这是我见过的 Claude 模型中最好的——但为此收了我 3.30 美元。
Claude has caught up a little bit. Claude Fable 5 gave me an _excellent_ pelican riding a bicycle—the best I've seen from a Claude model—but did charge me $3.30 for it.
Opus 5.5 思考了 128,000 个 token,然后放弃了!它在给出回答之前就用完了 token。
Opus 5.5 thought for 128,000 tokens and then gave up! It ran out of tokens before it got to the response.
回到 Deep Blue。今年一直让我困惑的是,为什么我的工作感觉更难了?
Getting back to Deep Blue. Something that's been puzzling me this year is why does my job feel harder?
我有了这些智能体,它们能为我完成所有这些事情,然而我从未如此努力地工作过,也从未对工作如此投入智力。
I've got these agents that can do all of this stuff for me, and yet I've never worked so hard, I've never been so intellectually engaged with my work.
部分原因是我承担的任务更具雄心,但也是因为所有简单的事情都由智能体处理了。如果事情简单,智能体就会去做。留给我的所有事情都是困难的。
Partly this is because I'm being a lot more ambitious with what I take on, but it's also because all of the easy stuff is handled for me. If it's easy, the agent will do it. Everything that's left for me is difficult.
今天早上,我听到了三届环法自行车赛冠军格雷格·莱蒙德的这句话:
This morning I heard this quote from three-time Tour de France champion, Greg LeMond:
我认为,这正是我们作为软件工程师,如今借助编码智能体所经历的情况。
I think that's exactly what's happening to us now as software engineers with coding agents.
最后一件收尾的事。我知道你们迫切想知道鸮鹦鹉繁殖季的最新情况。
One last closing thing. I know you're desperate for an update on Kākāpō breeding season.
我们达到了恢复期以来的新高——325 只鸟!
We've reached a recovery-era high of 325 birds!
已有 89 只新雏鸟存活至今。这是很长一段时间以来最好的繁殖年。
89 new chicks have made it to this point. This is the best breeding year in a very long time.
我听说 Claude Opus 5.5 现在能做像素艺术了。Claude 没有图像生成器,但它非常擅长用 JavaScript 绘制动画像素。
I heard that Claude Opus 5.5 can now do pixel art. Claude doesn't have an image generator, but it's very good at using JavaScript to draw animated pixels.
于是我让它给我做了一个鸮鹦鹉舞会。我觉得这是对今年最重要新闻的一个很好的庆祝。
So I had it make me a Kākāpō dance party. I think this is a good celebration of the most important news of this year.