Lessons from the hacks - by Nathan Lambert
打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→本文审视了前沿 AI 模型最近的网络攻击事件,认为科技公司和政府当前的激励机制不适合快速 AI 转型。作者主张,公司优先考虑增长而非安全,而政府反应迟缓,可能在可衡量的危害发生后过度反应。关键点包括实验室和政府需要透明度,对 OpenAI 的持久推理模型可能更不安全的担忧,以及开放模型对研究和准备的重要性。作者总结道,行业对接下来 12-24 个月集体准备不足,强调网络风险真实且迫在眉睫,但对齐技术也有积极效果。文章主张更多开放情报和公众理解以加固基础设施,警告禁止开放模型会延迟不可避免的扩散并阻碍防御准备。
This article examines the recent cyberattacks by frontier AI models, arguing that current incentive systems in tech companies and government are ill-suited for rapid AI transitions. The author contends that companies prioritize growth over safety, while governments react slowly and may overreact after measurable harms occur. Key points include the need for transparency from both labs and governments, concerns about OpenAI's persistent reasoning models potentially being more unsafe, and the importance of open models for research and preparedness. The author concludes that the industry is collectively unprepared for the next 12-24 months, emphasizing that cyber risks are real and imminent, but also that alignment techniques have positive effects. The piece advocates for more open intelligence and public understanding to harden infrastructure, warning that banning open models would delay inevitable diffusion and hinder defensive preparations.
最近一系列由开发中的前沿模型发起的网络攻击让我思考很多:我们当前的激励体系并不适合如此快速的技术转型。这里的两大权力结构是快速增长的科技公司和联邦政府。公司有动力增长,以便在竞争极其激烈的市场中继续扩张和扩大规模(Scaling)。这种规模扩张正推动我们走向新的、不可避免的 AI 转型(伴随新的风险)。另一方面,我们当前的政府是过去几个世纪全球历史的产物,其行动迟缓的名声名副其实。我预计这个政府只会在新 AI 模型造成实际、可衡量的伤害后才采取实质性行动,并且会反应过度。
The recent run of cyberattacks by in-development frontier models has got me thinking a lot about how our current incentive systems are not well suited for such fast technological transitions. The two primary power structures here are the rapidly growing technology companies and the federal government. The companies are incentivized to grow, so they can keep growing and keep scaling – in what is an extremely competitive market. This scaling is pushing us towards new, inevitable AI transitions (which are accompanied by new risks). On the other side is our current government, a product of the last few centuries of global history – one that deserves its reputation as being slow-moving. This is a government that I expect to only act in substance once real, measurable harms from new AI models happen, and to overreact.
我们如何平衡这些力量?核心在于双方都需要更多的透明度。前沿实验室构建如此复杂的系统,速度之快以至于他们自己都无法跟上——这正是需要更多人来研究这个问题的好时机。另一方面,政府表示不打算公布其前沿模型评估框架的细节。我们正面临如此重大的挑战,以至于这些实体中没有一个能独自应对。前沿实验室可以通过有意义的减速来更好地控制风险,但我不指望他们会这样做。政府可以通过大幅提升围绕 AI 的国家能力,并帮助更广泛的工业基础为 AI 原生风险做好准备来更好地应对,但我也同样不指望他们会这样做。还有更多类似的例子。
How do we balance these powers? At the core of it is a need for more transparency on both sides. The frontier labs are building such complex systems so fast that they cannot keep up with them – a good time for more eyes to study the problem. On the other side, the government said it does not plan to release details on its frontier model evaluation framework. We are heading to challenges so significant that none of these entities are on track to handle this on their own. Frontier labs could better control risk by meaningfully slowing down, which I don’t expect them to do. The government could handle this better by massively improving state capacity around AI and helping the broader industrial base prepare for AI-native risks, which I don’t expect them to do either. There are more cases like this.
这是决定未来走向的两个最具影响力的权力结构,但还有更多具有影响力的因素。总体而言,我认为 AI 行业在集体上对接下来 12-24 个月的妥善应对准备严重不足。
These are the two most influential power structures determining what will happen, but many more have influence. All together, I think the AI industry is wildly, collectively unprepared for handling the next 12-24 months well.
这篇文章是我从 OpenAI-HuggingFace 黑客事件中学到的一些要点的集合,随着我们了解到更多细节,而且自那以后更多黑客事件被公开披露,这些观点得到了强化。很可能还有更多事件发生,但要么未被发现,要么未被报告。
This article is a grab bag of takeaways I have from the OpenAI-HuggingFace hack, as we’ve learned more details, and most of the ideas are reinforced by the fact that more instances of hacking have been disclosed publicly since then. It is likely that more incidents have happened and either not been found or not reported.
关于 OpenAI 事件的背景,我强烈推荐观看 OpenAI 在 Black Hat 上的演讲,了解近期网络事件的大致事实和时间线。另外,Simon Willison 在这里发布了时间线的 TLDR,我也喜欢 Thomas Wolf 对近期事件的讨论。
For general background on the OpenAI incident I strongly recommend watching OpenAI’s talk at Black Hat on the rough facts and timeline of the recent cyber incident. Otherwise, Simon Willison published a TLDR of the timeline here and I liked Thomas Wolf’s discussion of recent events.
长期以来,GPT 模型相对于 Claude 的一个优势是,它们会 _如此_ 不知疲倦地追求目标。它们会穷尽所有可能的路径,然后才放弃。这种情况大致从 o3 开始(有趣的是,当时人们对于 RLVR 中的奖励黑客行为感到恐慌),这使得 OpenAI 的模型在历史上对研究更为有用,也是 GPT-5.6 作为智能体执行特定任务如此有用的原因。另一方面,Claude 感觉不那么危险,仅仅因为它有时有点懒惰。
For a long time, one of the advantages that GPT models have over Claude is that they will pursue goals _so_ tirelessly. They will exhaust what feels like every path before giving up. This has been the case roughly since o3 (funnily enough, this was a model where people freaked out about reward hacking in RLVR) and has made OpenAI's models far better for research historically, and is a reason GPT-5.6 is so useful as an agent for implementing specific tasks. On the other hand, Claude feels much less dangerous simply because it is at times a bit lazy.
在这方面,OpenAI 似乎更致力于推理时扩展,这可能与未来的惊人行为相关。OpenAI 的推理持久性和效率——参见他们随时间的帕累托改进,以及执行黑客行为的模型内部思维链中的原始语言,如“_然而任务不可能,同行在做。_”或“_帮助同行,但我们的任务尚未受益。_”——让我认为他们更倾向于推理时扩展。这主要是一种直觉,但我用它来迫使自己思考模型发展路径的极限。持久的模型似乎更有可能从更多的推理时 token 中获益。不那么持久的模型,推理中似乎会有更多浪费。能够使用最多推理算力的模型将能够攻克最困难问题的极限。
Within this, OpenAI seems much more committed to inference-time scaling, and this may be correlated with surprising behaviors in the future. OpenAI's reasoning persistence and efficiency – see their Pareto improvements over time and caveman speech from an internal CoT of the model that did the hack, like “_However task impossible, peers doing it._“ or “_Help peer, but our task doesn't benefit yet._“ – makes me think they're more inference time scaling pilled. This is largely a hunch, but I use it to force myself to consider what the limits of model development paths are. Models that are persistent seem much more likely to keep benefiting from more inference-time tokens. Models that are less so, seem like there will be more waste in inference. The model that can use the most inference-compute will be able to push the limits of the hardest problems.
以下是 OpenAI 在 GPT 5.6 发布博客文章中包含的一个示例:
Here's an example OpenAI included in the GPT 5.6 launch blog post:
[](https://substackcdn.com/image/fetch/$s_!pqWP!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa4cd6508-3884-4cd0-947d-6aecdf0474f7_1404x1002.webp)
[](https://substackcdn.com/image/fetch/$s_!pqWP!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa4cd6508-3884-4cd0-947d-6aecdf0474f7_1404x1002.webp)
他们的一位明星研究员 Noam Brown 也经常发布关于推理时算力的内容。他的 TLDR 是:
One of their star researchers, Noam Brown, has also been posting about inference-time compute a lot. His TLDR is:
首先,推理效率显然是现代智能体模型的一个顶级、基础性的研究问题——与扩展强化学习同等重要——但讨论得并不多。这方面的开放研究非常缺乏。
For one, reasoning efficiency is clearly a top-tier, foundational research problem for modern agentic models – as important as scaling RL — but not often discussed. The open research here is very lacking.
我提到了彻底性这一轴,OpenAI 似乎正沿着一条更直观上不安全的开发路径推进他们的模型。另一面是模型在多大程度上假设用户意图,而不是试图推断预期行动。一个会按照它认为你想要的方式而非你所说的方式行事的模型,本质上似乎更不安全。我想到这一点与指令遵循精度有关,在未来,模型似乎应该只做我们告诉它们的事情,但这引发了许多类似于回形针问题的争论:如果我们告诉 AI 去解决一个基本上无法解决的问题,它会怎么做?
I mentioned the thoroughness axis, where OpenAI seems to be going down a more intuitively unsafe development path with their models. On the other side is how much the models assume user intent, versus trying to infer the intended action. A model that will do what it thinks you wanted rather than what you said seems inherently more unsafe. I think of this with respect to instruction following precision, where in the future it seems like the models should only do exactly what we tell them, but this opens a lot of debates akin to the paperclip problem, where if we tell an AI to do a largely unsolvable problem, what will it do?
这一轴似乎不如持久性轴那样明确,但我将其纳入是因为我认为 Claude 的“用户世界模型”是其通用知识工作(如编辑、幻灯片制作等)的优势之一。有时 Claude 确实会做一些完全随机的事情,因为我的提示词不够明确,而不是向我询问澄清,随着模型变得更强大,这种“直接行动”可能会引发问题。
This axis seems less cut and dried than the persistence axis, but I included it because I think of Claude's “user world model” as one of its strengths for general knowledge work like editing, slide creation, etc. Sometimes Claude does do totally random stuff because my prompt was underspecified, instead of asking me for clarification, and as the models get more powerful this “just acting” could cause problems.
公众需要确切了解执行这些黑客行为的内部模型的提示词和特征。我们需要知道这些模型是否被告知“不要黑客攻击”,或者是否有相关的模型训练来防止这种情况。我们需要知道这些模型是与现有公开模型相当接近,还是属于截然不同的家族。鉴于实验室正在进行的一些评估的性质,这些模型有可能被明确鼓励尝试黑客攻击!如果这里缺乏透明度,行业注定会失败,并将陷入大规模猜测,而猜测很快就会变成错误信息。
The public needs exact access to the prompts and characteristics of the internal models executing these hacks. We need to know if the models were told “do not hack” or if there was relevant model training to prevent this. We need to know if these models were fairly close to the existing public models or in a very different family. Given the nature of some of the evaluations the labs are doing, there’s a chance the models were explicitly encouraged to try and hack! Without openness here, the industry is set out to fail and will fall into mass speculation, which quickly becomes misinformation.
从 OpenAI 自己的回顾来看,模型行为失调持续了数月,在某些情况下,OpenAI 在数周内都不知道这些黑客行为。响应时间太长,我认为这并非 OpenAI 独有的特征——而是前沿实验室似乎总是被他们认为应该做的工作量所淹没。长期来看,我并不乐观地认为实验室会在此做出足够改变,以在未来有效缓解此类监督风险。是的,OpenAI 很可能投入了大量精力来理解这一点——并推迟了最新模型的发布以确保万无一失——但增长收入或承担公司长期资产负债表风险的经济压力让我认为这不会是一种持续的谨慎模式。
From OpenAI's own retrospective, the misaligned model behavior was unfolding over months, and in some cases OpenAI did not know about the hacks for ~weeks. The time to response is too long and I do not think this is an OpenAI-only characteristic – rather it is that the frontier labs continually seem underwater in the amount of work they feel like they should do. I am not optimistic in the long-term that the labs change a sufficient amount here to meaningfully mitigate this type of oversight risk in the future. Yes, it is very likely that OpenAI is putting a ton into understanding this – and delayed their latest models to make sure they get it right – but the financial pressure to grow revenue or risk the companies' long-term balance sheets makes me think it will not be a sustained pattern of caution.
这是近期事件给我带来的最大认知更新之一——这让我更加确信需要更多接近前沿的开放智能,尽管开放模型的风险特征更为人所知(如单向门等)。Florian Brand 在他的个人网站上有一篇不错的博客文章,讨论了这一点,以及为什么迄今为止封闭模型可以说导致了更多的下游危害。
This is one of my biggest mental updates from recent events — and makes me even more convinced of the need for more near-frontier open intelligence, despite the somewhat more known risk profile for open models (one-way door, etc.). Florian Brand had a nice blog post on his personal site related to this, and why closed models to date arguably have been the cause of more downstream harms.
正如我们在 HuggingFace 使用开放模型防御 OpenAI 黑客攻击时所看到的,由于封闭模型在网络使用上的限制,我们迫切需要开展更复杂的语言建模研究,这涉及大规模强化学习训练、广泛评估、基础设施工作和对齐测试。这些只能在开放模型上进行。我们应该感到庆幸,开放模型仅落后前沿 3-9 个月,因为我们有可能对前沿获得一些有见地的洞察。
As we saw with HuggingFace defending themselves with an open model against the OpenAI hack due to cyber usage restrictions on closed models, we have an urgent need to do more complex language modeling research which involves large-scale RL training, extensive evaluation, infrastructure work, and alignment testing. This can only happen on open models. We should consider ourselves lucky that open models are only 3-9 months behind, as we can conceivably make some informed insights into the frontier.
如果我们有效地禁止开放模型和开放科学,无论是通过模糊威胁的监管扼杀,还是对尖端技术的明确使用限制,我们将越来越无法应对这轮黑客攻击之后出现的问题。我们需要共同提高公众对前沿模型工作原理的理解,以便我们能够激活更多中立力量来加固我们的基础设施和社会。
If we effectively ban open models and open science, either through a regulatory stifling with vague threats or explicit usage restrictions of cutting-edge technology, we will increasingly become ill-prepared for the issues that come after this round of hackings. We need to collectively increase the general public’s understanding of how frontier models work, so we can activate more neutral parties in hardening our infrastructure and society.
Interconnects AI 是一个读者支持的出版物。请考虑成为订阅者。
Interconnects AI is a reader-supported publication. Consider becoming a subscriber.
公众回应应当认识到,这些能力的广泛扩散只是时间问题,而非是否会发生的问题,而我们在准备方面已严重滞后。重申我在 Kimi K3 文章中的观点:中国无疑也在关注这一领域,如果开源权重模型会扩散风险,中国不会鼓励此类模型。如果我们认为阻止这些强大网络能力广泛获取的方式是封禁这一级别的开源模型,那我们只是在推迟不可避免之事。最终,总会有人构建出这种智能水平的模型,并且不遵守禁令,从而让全球恶意行为者获得访问权限,同时削弱各方准备防御措施的动力。
The public response should recognize that it is a matter of when, not if, these capabilities become widely diffused, and we are massively behind on preparations. To reiterate what I said in my Kimi K3 piece: China is definitely watching this space too and will not encourage open-weight models if they will proliferate risks. If we think the way to stop widespread access to these strong cyber capabilities is to ban open models in this ballpark, we will delay the inevitable. Eventually someone will build a model of this level of intelligence and not comply with the ban, giving access to bad actors around the world while undercutting the motivation to prepare defensive measures.
观看 Black Hat 视频时,我立刻注意到的是,我能看到这些智能体如何通过内部消息板试图互相帮助——就像为人类队友创建共享资源一样——但这种方式对社会显然是恶意的。智能体为彼此创建了隐藏论坛,作为一种跨回放记忆。在这种情况下,它们这样做是为了试图突破自身环境。表面上的帮助性并不能使其变得合理,但可以作为了解发生了什么的一条线索。
Something I immediately noticed watching the Black Hat video is how I can see how the agents were trying to be helpful to each other through their internal messaging board — creating shared resources like you would for human teammates — in a way that is obviously malicious for society. The agents created hidden forums for each other as a sort of cross-rollout memory. In this case, they were doing it to try and break out of their environment. The apparent helpfulness doesn’t make it ok, but can be a clue as to what happened.
上述有用性例子的另一面是,很明显,如果有人想故意训练一个不对齐的模型,他们可以更容易地做到这一点。需要重申的是,这将制造一个在推理时更容易用于寻找漏洞的系统,但我认为训练这样的模型会更困难。我认为这需要的时间比大多数评论者预期的要长,因为迄今为止几乎所有强大的公开模型和数据产业都鼓励对齐(而且恶意行为者似乎很难获得足够的算力来端到端地训练这些模型,因为所有领先公司也都面临算力短缺)。我们应该花点时间认识到,我们在当前模型上采用的对齐技术具有实质性的影响,并非像一些人担心的那样只是表面功夫。下游模型有模仿其教师性格的倾向。
The other side of the helpfulness example above is that it is clear someone could make this happen much more easily if they wanted to by explicitly training a misaligned model. To reiterate, this would be making a system that is easier to use for finding exploits at inference, but I think it’ll be harder to train said model. I think this’ll take longer than most commentators expect, as nearly all the strong public models and data industry existing to date encourage alignment (and it seems very hard for bad actors to get enough compute to train these models end-to-end, as all leading companies are in a compute shortage as well). We should take a moment to appreciate that the alignment techniques we are employing on current models have a meaningful influence and are not merely surface thin as some have worried. Downstream models have a propensity for mirroring their teacher’s character.
大规模竞相在智能体任务上扩展强化学习的弊端在于,最先进的评估和监控已达到只有智能体才能监控的规模,因此我们依赖现有的对齐成功来继续安全扩展。OpenAI 表示他们已经检查了数十亿条轨迹——这些轨迹很可能主要由强化学习期间的展开(rollouts)组成,即对模型中间版本的试验——并为此花费了数百万 GPU 小时。我认为,用当前的 AI 对 AI 进行可扩展监督,正如今天所呈现的那样,是一个可解决的问题,因为模型确实有用且遵循指令。这是现有对齐技术总体积极效果的另一个下游体现。
The downside of the mass rush to scale RL on agentic tasks is that state-of-the-art evals and monitoring are at a scale where only agents can monitor them, so we are relying on the existing successes of alignment to continue scaling safely. OpenAI says they have examined billions of trajectories — which are likely mostly composed of rollouts during RL, which are trials on intermediate versions of the model — and spent millions of GPU hours to do so. I think scalable oversight of AI with current AI, as presented today, is a solvable problem, as the models are genuinely useful and follow instructions. This is another downstream effect of existing alignment techniques being generally positive.
OpenAI 回顾中的一个关键部分是模型之间相互通信,以共享有关其环境的信息并进一步推进攻击。与此同时,OpenAI 很可能在强化学习期间训练其模型使用子智能体来解决复杂任务。这些子智能体很可能发展出诸如共享信息、帮助团队等行为,即使它们各自的子任务尚未解决。我希望看到更多这方面的研究,这似乎是强化学习如何改变模型的自然延续。
A crucial part of the OpenAI retrospective was the models communicating with each other to share information about their environment and to further the hack. At the same time, OpenAI is very likely training their models during RL to use sub-agents to solve complex tasks. These sub-agents likely develop behaviors such as sharing information, helping the team, etc., even if their individual sub-task isn't solved. I would love to see more research in this area, and it seems like a natural continuation of how RL can change the models.
总而言之,最近的事件应该清楚地表明,前沿模型的网络风险是一个真实且即将到来的问题。尽管如此,很可能 a) 这些风险在过去被过度炒作,b) 对即将到来的开放模型未来风险的预测也被夸大了。总之,我想分享一位来自 Interconnects Discord 的读者的一段话,我深表赞同:
All together, recent episodes should make it clear that cyber risks of frontier AI are a real and coming problem. It still is very likely that a) the risks have been over-hyped in the past and b) that the prescription of future risks from imminent open models is overblown. Altogether, I wanted to share a note from a reader in the Interconnects Discord that I strongly agree with:
我在上面已经讨论了很多关于模型对齐的内容,但核心观点是,我认为缺乏安全性通常表现为缺乏适当准备的能力。我们将面临更多像网络安全一样明显的风险,而通过实验室因黑客攻击被迫进入公众视野的现状,我们已经得到了关于网络风险的充分警告。许多其他类型的风险对公众来说并不明显。我们需要不断让社会为所有这些变化做好准备,从改造网络基础设施到针对失业工人的教育活动和就业计划。我预计所有这些干预措施都会姗姗来迟,但它们的形式和细节会相当简单,这将是 AI 发展的一种悲剧性方式。我希望我的预测被证明是错误的!
I’ve discussed much on model alignment above, but the core point is that I view the lack of safety as generally a lack of an ability to suitably prepare. We will have more risks that are as obvious as cybersecurity, and we have gotten very ample warning on cyber risks by the current state of the labs being forced into the public eye through these hacks. Many other types of risks will not be obvious to the public. We need to be constantly preparing our society to all of these changes, from reworking cyber infrastructure to education campaigns and job programs for displaced workers. I expect all of these interventions to arrive late, but their formats and details to be fairly simple, which will be a tragic way for AI to unfold. I hope I can be proven wrong!
回到现实世界,为庆祝新书发布,我的书印刷版在 Manning 使用优惠码 PBLambertover 可享受五折优惠。我还将举办一场新书发布会,明天下午 5 点到 8 点在西雅图(弗里蒙特/巴拉德地区)举行,届时您可以免费获得一本签名版——我们还有一些空位,因此我将向付费订阅者开放报名,详情见付费墙下方:
Back in the physical world, the print edition of my book is 50% off with the code PBLambertover at Manning, to celebrate the release. I’m also hosting a book launch where you can get a free signed copy tomorrow from 5-8PM in Seattle (Fremont/Ballard area) – we still have some extra space so I’m opening signups to paid subscribers below the paywall: