The Urgency of Interpretability
打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→机制可解释性的简史
* A Brief History of Mechanistic Interpretability
* 机制可解释性简史
* A Brief History of Mechanistic Interpretability
在我从事人工智能研究的这十年里,我目睹了它从一个微小的学术领域成长为世界上最重要的经济和地缘政治问题。在这段时间里,我学到的最重要的一课或许是:底层技术的进步是不可阻挡的,由过于强大的力量驱动,但它的发生方式——构建的顺序、我们选择的应用以及向社会推广的细节——是完全可以改变的,并且通过这样做可以产生巨大的积极影响。我们无法阻止这辆巴士,但我们可以驾驶它。过去,我曾写过关于以对世界积极的方式部署人工智能的重要性,以及确保民主国家在专制国家之前构建和掌握这项技术。在过去的几个月里,我越来越关注另一个驾驶巴士的机会:最近的一些进展开辟了一个诱人的可能性,即我们可以在模型达到压倒性力量之前成功实现可解释性——也就是说,理解人工智能系统的内部运作。
In the decade that I have been working on AI, I’ve watched it grow from a tiny academic field to arguably the most important economic and geopolitical issue in the world. In all that time, perhaps the most important lesson I’ve learned is this: the progress of the underlying technology is inexorable, driven by forces too powerful to stop, but the _way_ in which it happens—the order in which things are built, the applications we choose, and the details of how it is rolled out to society—are eminently possible to change, and it’s possible to have great positive impact by doing so. We can’t _stop_ the bus, but we can _steer_ it. In the past I’ve written about the importance of deploying AI in a way that is positive for the world, and of ensuring that democracies build and wield the technology before autocracies do. Over the last few months, I have become increasingly focused on an additional opportunity for steering the bus: the tantalizing possibility, opened up by some recent advances, that we could succeed at _interpretability_—that is, in understanding the inner workings of AI systems—_before_ models reach an overwhelming level of power.
外界的人常常感到惊讶和警觉,得知我们并不了解自己创造的人工智能系统是如何工作的。他们的担忧是有道理的:这种缺乏理解在技术史上基本上是前所未有的。几年来,我们(Anthropic 以及整个领域)一直在努力解决这个问题,试图创建一种类似于高度精确和准确的 MRI 的东西,以完全揭示人工智能模型的内部运作。这个目标常常感觉非常遥远,但最近的几项突破让我相信,我们现在走在正确的轨道上,并且有真正的成功机会。
People outside the field are often surprised and alarmed to learn that we do not understand how our own AI creations work. They are right to be concerned: this lack of understanding is essentially unprecedented in the history of technology. For several years, we (both Anthropic and the field at large) have been trying to solve this problem, to create the analogue of a highly precise and accurate MRI that would fully reveal the inner workings of an AI model. This goal has often felt very distant, but multiple recentbreakthroughs have convinced me that we are now on the right track and have a real chance of success.
与此同时,整个人工智能领域比我们在可解释性方面的努力更先进,并且自身发展非常迅速。因此,如果我们希望可解释性及时成熟并发挥作用,就必须迅速行动。这篇文章阐述了可解释性的理由:它是什么,为什么拥有它会让人工智能发展得更好,以及我们所有人可以做些什么来帮助它赢得这场竞赛。
At the same time, the field of AI as a whole is further ahead than our efforts at interpretability, and is itself advancing very quickly. We therefore must move fast if we want interpretability to mature in time to matter. This post makes the case for interpretability: what it is, why AI will go better if we have it, and what all of us can do to help it win the race.
现代生成式 AI 系统的不透明性从根本上不同于传统软件。如果一个普通软件程序做了某事——例如,视频游戏中的角色说了一句台词,或者我的外卖应用允许我给司机小费——它做这些事是因为人类专门编程设定了它们。生成式 AI 则完全不同。当一个生成式 AI 系统做某事时,比如总结一份财务文档,我们完全不知道,在具体或精确的层面上,它为什么做出这样的选择——为什么选择某些词而不是其他词,或者为什么它偶尔会犯错尽管通常很准确。正如我的朋友兼联合创始人 Chris Olah 喜欢说的那样,生成式 AI 系统是“生长”出来的,而不是“构建”出来的——它们的内部机制是“涌现”的,而非直接设计的。这有点像种植植物或培养细菌群落:我们设定了引导和塑造生长的高层条件。
Modern generative AI systems are opaque in a way that fundamentally differs from traditional software. If an ordinary software program does something—for example, a character in a video game says a line of dialogue, or my food delivery app allows me to tip my driver—it does those things because a human specifically programmed them in. Generative AI is _not like that at all_. When a generative AI system does something, like summarize a financial document, we have no idea, at a specific or precise level, why it makes the choices it does—why it chooses certain words over others, or why it occasionally makes a mistake despite usually being accurate. As my friend and co-founder Chris Olah is fond of saying, generative AI systems are _grown_ more than they are _built_—their internal mechanisms are “emergent” rather than directly designed. It’s a bit like growing a plant or a bacterial colony: we set the high-level conditions that direct and shape growth1
1 以植物为例,这些条件包括水、阳光、引导其生长方向的棚架、选择植物种类等。这些因素大致决定了植物的生长方向,但其确切的形状和生长模式无法预测,甚至在生长后也难以解释。对于 AI 系统,我们可以设定基本架构(通常是 Transformer 的某种变体)、它们接收的数据的大致类型以及用于训练它们的高层算法,但模型的实际认知机制是从这些成分中有机涌现出来的,我们对它们的理解很有限。事实上,在自然和人工世界中都有许多例子,我们在原理层面理解(有时甚至控制)这些系统,但无法了解其细节:经济、雪花、元胞自动机、人类进化、人类大脑发育等等。
1 In the case of a plant, this would be water, sunlight, a trellis pointing them in a certain direction, choosing the species of plant, etc. These things dictate _roughly_ where the plant grows, but its exact shape and growth pattern are impossible to predict, and hard to explain even after they’ve grown. In the case of AI systems, we can set the basic architecture (usually some variant of the Transformer), the broad type of data they receive, and the high-level algorithm used to train them, but the model’s actual cognitive mechanisms emerge organically from these ingredients, and our understanding of them is poor. In fact, there are many examples, in both the natural and artificial worlds, of systems we understand (and sometimes control) at the level of principles but not in detail: economies, snowflakes, cellular automata, human evolution, human brain development, and so on.
但涌现出的确切结构是不可预测的,且难以理解或解释。观察这些系统内部,我们看到的是由数十亿个数字组成的巨大矩阵。它们以某种方式计算着重要的认知任务,但具体如何做到并不明显。
, but the exact structure which emerges is unpredictable and difficult to understand or explain. Looking inside these systems, what we see are vast matrices of billions of numbers. These are _somehow_ computing important cognitive tasks, but exactly how they do so isn’t obvious.
许多与生成式 AI 相关的风险和担忧最终都是这种不透明性的后果,如果模型是可解释的,这些问题会容易处理得多。例如,AI 研究者经常担心不对齐的系统可能采取其创造者未意图的有害行为。我们无法理解模型的内部机制,意味着我们无法有意义地预测此类行为,因此难以排除它们;事实上,模型确实表现出意外的涌现行为,尽管尚未达到令人严重担忧的程度。更微妙的是,同样的不透明性使得很难找到确凿证据来支持这些风险在大规模上的存在,从而难以凝聚力量去应对它们——实际上,也难以确切知道它们有多危险。
Many of the risks and worries associated with generative AI are ultimately consequences of this opacity, and would be much easier to address if the models were interpretable. For example, AI researchers often worry about misaligned systems that could take harmful actions not intended by their creators. Our inability to understand models’ internal mechanisms means that we cannot meaningfully predict such behaviors, and therefore struggle to rule them out; indeed, models _do_ exhibit unexpected emergent behaviors, though none that have yet risen to major levels of concern. More subtly, the same opacity makes it hard to find definitive evidence _supporting_ the existence of these risks at a large scale, making it hard to rally support for addressing them—and indeed, hard to know for sure how dangerous they are.
要解决这些对齐风险的严重性,我们必须比今天更清晰地看清 AI 模型的内部。例如,一个主要的担忧是 AI 的欺骗或权力寻求。AI 训练的性质使得 AI 系统有可能自行发展出欺骗人类的能力和寻求权力的倾向,这是普通确定性软件永远不会做到的;这种涌现性质也使得检测和缓解此类发展变得困难。
To address the severity of these alignment risks, we will have to see inside AI models much more clearly than we can today. For example, one major concern is AI deception or power-seeking. The nature of AI training makes it possible that AI systems will develop, on their own, an ability to deceive humans and an inclination to seek power in a way that ordinary deterministic software never will; this emergent nature also makes it difficult to detect and mitigate such developments2
2 当然,你可以尝试通过简单地与模型交互来检测这些风险,我们在实践中也这样做。但由于欺骗正是我们试图发现的行为,外部行为并不可靠。这有点像试图通过询问某人是否是恐怖分子来判断他是否是恐怖分子——并非毫无用处,你可以从他们的回答和言语中了解一些东西,但显然非常不可靠。
2 You can of course try to detect these risks by simply interacting with the models, and we do this in practice. But because deceit is precisely the behavior we’re trying to find, external behavior is not reliable. It’s a bit like trying to determine if someone is a terrorist by asking them if they are a terrorist—not necessarily useless, and you can learn things by how they answer and what they say, but very obviously unreliable.
但同样,我们从未在真正的现实场景中看到过欺骗和权力寻求的可靠证据,
. But by the same token, we’ve never seen any solid evidence in truly real-world scenarios of deception and power-seeking3
3 我可能会在未来的文章中更详细地描述这一点,但确实有很多实验(其中许多由 Anthropic 完成)表明,当训练以某种人为方式引导时,模型可以在特定情况下撒谎或欺骗。也有证据表明存在一些模糊地类似于“考试作弊”的现实行为,尽管它更多是退化性的而非危险或有害的。真正缺乏的是以更自然的方式涌现出危险行为的证据,或者为了获得对世界的权力而撒谎和欺骗的“普遍倾向”或“普遍意图”。在后者这一点上,看清模型内部可能大有帮助。
3 I’ll probably describe this in more detail in a future essay, but there _are_ a lot of experiments (many of which were done by Anthropic) showing that models can lie or deceive under certain circumstances when their training is guided in a somewhat artificial way. There is also evidence of real-world behavior that looks vaguely like “cheating on the test”, though it’s more degenerate than it is dangerous or harmful. What there _isn’t_ is evidence of dangerous behaviors emerging in a more naturalistic way, or of a _general tendency_ or _general intent_ to lie and deceive for the purposes of gaining power over the world. It is the latter point where seeing inside the models could help a lot.
因为我们无法“当场抓住模型”正在思考权力欲强、欺骗性的想法。我们剩下的只是模糊的理论论证,即欺骗或权力寻求可能在训练过程中有涌现的激励,有些人觉得这完全令人信服,而另一些人则觉得可笑且毫无说服力。老实说,我对这两种反应都能理解,这或许能说明为什么关于这一风险的辩论变得如此两极分化。
because we can’t “catch the models red-handed” thinking power-hungry, deceitful thoughts. What we’re left with is vague theoretical arguments that deceit or power-seeking might have the incentive to emerge during the training process, which some people find thoroughly compelling and others laughably unconvincing. Honestly I can sympathize with both reactions, and this might be a clue as to why the debate over this risk has become so polarized.
类似地,对 AI 模型被滥用的担忧——例如,它们可能帮助恶意用户制造生物或网络武器,其方式超出了当今互联网上可找到的信息——是基于
Similarly, worries about misuse of AI models—for example, that they might help malicious users to produce biological or cyber weapons, in ways that go beyond the information that can be found on today’s internet—are based4
4 至少在 API 服务模型的情况下是这样。开放权重模型带来了额外的危险,因为防护措施可以被轻易剥离。
4 At least in the case of API-served models. Open-weights models present additional dangers in that guardrails can be simply stripped away.
这样一种观点:很难可靠地阻止模型知道危险信息或泄露它们所知的内容。我们可以给模型加上过滤器,但有大量可能的方法来“越狱”或欺骗模型,而发现越狱存在的唯一方法是经验性地找到它。相反,如果能够查看模型内部,我们或许能够系统地阻止所有越狱,并描述模型拥有哪些危险知识。
on the idea that it is very difficult to reliably prevent the models from knowing dangerous information or from divulging what they know. We can put filters on the models, but there are a huge number of possible ways to “jailbreak” or trick the model, and the only way to discover the existence of a jailbreak is to find it empirically. If instead it were possible to look inside models, we might be able to systematically block all jailbreaks, and also to characterize what dangerous knowledge the models have.
AI 系统的不透明性也意味着它们根本无法在许多应用中使用,例如高风险的金融或安全关键场景,因为我们无法完全设定其行为边界,而少量错误可能造成极大危害。更好的可解释性可以极大地提高我们设定可能错误范围的能力。事实上,对于某些应用,无法查看模型内部在法律上直接阻碍了其采用——例如在抵押贷款评估中,决策依法需要可解释。类似地,AI 在科学领域取得了巨大进步,包括改进 DNA 和蛋白质序列数据的预测,但以这种方式预测的模式和结构往往难以被人类理解,也无法带来生物学洞见。过去几个月的一些研究论文已经明确表明,可解释性可以帮助我们理解这些模式。
AI systems’ opacity also means that they are simply not used in many applications, such as high-stakes financial or safety-critical settings, because we can’t fully set the limits on their behavior, and a small number of mistakes could be very harmful. Better interpretability could greatly improve our ability to set bounds on the range of possible errors. In fact, for some applications, the fact that we can’t see inside the models is literally a legal blocker to their adoption—for example in mortgage assessments where decisions are legally required to be explainable. Similarly, AI has made great strides in science, including improving the prediction of DNA and protein sequence data, but the patterns and structures predicted in this way are often difficult for humans to understand, and don’t impart biological insight. Some research papers from the last few months have made it clear that interpretability canhelp us understand these patterns.
还有更多其他因不透明性导致的后果,例如它阻碍了我们判断 AI 系统是否(或可能有一天)具有感知能力,并可能值得拥有重要权利。这是一个足够复杂的话题,我不会详细展开,但我认为它在未来会很重要。
There are other more exotic consequences of opacity, such as that it inhibits our ability to judge whether AI systems are (or may someday be) sentient and may be deserving of important rights. This is a complex enough topic that I won’t get into it in detail, but I suspect it will be important in the future.5
5 非常简要地说,可解释性可能以两种方式与对 AI 感知和福利的关切相交织。首先,尽管心灵哲学是一个复杂且有争议的话题,但哲学家无疑会从对 AI 模型中实际发生情况的详细描述中受益。如果我们认为它们只是浅层的模式匹配器,那么它们似乎不太可能值得道德考虑。如果我们发现它们执行的计算与动物甚至人类的大脑相似,那可能是支持道德考虑的证据。其次,也许最重要的是,如果我们得出结论认为 AI 的道德“患者地位”足够合理以至于需要采取行动,那么可解释性将发挥关键作用。对 AI 的严肃道德考量不能相信它们的自我报告,因为我们可能无意中训练它们在自己不好时假装没事。在这种情况下,可解释性在确定 AI 的福祉方面将具有关键作用。(事实上,从这个角度来看,已经有一些略微令人担忧的迹象。)
5 Very briefly, there are two ways in which you might expect interpretability to intersect with concerns about AI sentience and welfare. Firstly, while philosophy of mind is a complex and contentious topic, philosophers will no doubt benefit from a detailed accounting of what actually is occurring in AI models. If we believe them to be superficial pattern-matchers, it seems unlikely they warrant moral consideration. If we find that the computation they perform is similar to the brains of animals, or even humans, that might be evidence in favor of moral consideration. Secondly, and perhaps most importantly, is the role interpretability would have if we ever concluded that the moral “patienthood” of AI models was plausible enough to warrant action. A serious moral accounting on AI can't trust their self-reports, since we might accidentally train them to pretend to be okay when they aren't. Interpretability would have a crucial role in determining the wellbeing of AIs in such a situation. (There are, in fact, already some mildly concerning signs from this perspective.)
基于上述所有原因,弄清楚模型在想什么以及它们如何运作似乎是一项极其重要的任务。几十年来,传统观点认为这是不可能的,模型是不可理解的“黑箱”。我无法完整讲述
For all of the reasons described above, figuring out what the models are thinking and how they operate seems like a task of overriding importance. The conventional wisdom for decades was that this was impossible, and that the models were inscrutable “black boxes”. I’m not going to be able to do justice6
例如,自 70 多年前神经网络发明以来,以某种方式分解和理解人工神经网络内部计算的想法可能就隐约存在,而理解神经网络为何以特定方式行为的各种努力也几乎同样历史悠久。但克里斯·奥拉的不同之处在于,他提出并认真追求一项全面努力,以理解模型所做的一切。
6 For example, the idea of somehow breaking down and understanding the computations happening inside artificial neural networks was probably around in a vague sense since neural networks were invented over 70 years ago, and various efforts to understand why a neural net behaved in a specific way have existed for nearly as long. But Chris was unusual in proposing _and_ seriously pursuing a comprehensive effort to understand _everything_ they do.
关于这一观念如何转变的完整故事,我的观点不可避免地受到我在谷歌、OpenAI 和 Anthropic 亲身经历的影响。但克里斯·奥拉是最早尝试真正系统性研究计划以打开黑箱并理解其所有组成部分的人之一,这一领域后来被称为机械可解释性。克里斯先在谷歌,后在 OpenAI 从事机械可解释性研究。当我们创立 Anthropic 时,我们决定将其作为新公司方向的核心部分,并且关键的是,将其聚焦于大语言模型。随着时间的推移,该领域不断发展,现在包括几家主要 AI 公司的团队,以及一些专注于可解释性的公司、非营利组织、学术界和独立研究人员。简要总结该领域迄今取得的成就,以及如果我们希望应用机械可解释性来解决上述关键风险,还有哪些工作要做,这将是有益的。
to the full story of how that changed, and my views are inevitably colored by what I saw personally at Google, OpenAI, and Anthropic. But Chris Olah was one of the first to attempt a truly systematic research program to open the black box and understand all its pieces, a field that has come to be known as _mechanistic interpretability_. Chris worked on mechanistic interpretability first at Google, and then at OpenAI. When we founded Anthropic, we decided to make it a central part of the new company’s direction and, crucially, focused it on LLM’s. Over time the field has grown and now includes teams at several of the major AI companies as well as a few interpretability-focused companies, nonprofits, academics, and independent researchers. It’s helpful to give a brief summary of what the field has accomplished so far, and what remains to be done if we want to apply mechanistic interpretability to address some of the key risks above.
机械可解释性的早期时代(2014-2020 年)专注于视觉模型,并能够识别模型内部一些代表人类可理解概念的神经元,例如“汽车检测器”或“轮子检测器”,类似于早期神经科学假设和研究,表明人脑具有对应特定人物或概念的神经元,通常被通俗地称为“詹妮弗·安妮斯顿神经元”(事实上,我们在 AI 模型中也发现了类似的神经元)。我们甚至能够发现这些神经元是如何连接的——例如,汽车检测器会寻找汽车下方激活的轮子检测器,并结合其他视觉信号来判断它看到的物体是否确实是汽车。
The early era of mechanistic interpretability (2014-2020) focused on vision models, and was able to identify some neurons inside the models that represented human-understandable concepts, such as a “car detector” or a “wheel detector”, similar to early neuroscience hypotheses and studies suggesting that the human brain has neurons corresponding to specific people or concepts, often popularized as the “Jennifer Aniston” neuron (and in fact, we found neurons much like those in AI models). We were even able to discover how these neurons are connected—for example, the car detector looks for wheel detectors firing below the car, and combines that with other visual signals to decide if the object it’s looking at is indeed a car.
当克里斯和我离开创立 Anthropic 时,我们决定将可解释性应用于新兴的语言领域,并在 2021 年开发了一些必要的基本数学基础和软件基础设施。我们立即在模型中发现了执行语言理解所必需的基本机制:复制和顺序模式匹配。我们还发现了一些可解释的单个神经元,类似于我们在视觉模型中发现的那样,它们代表各种单词和概念。然而,我们很快发现,虽然某些神经元是直接可解释的,但绝大多数神经元是许多不同单词和概念的混乱拼凑。我们将这种现象称为叠加,
When Chris and I left to start Anthropic, we decided to apply interpretability to the emerging area of language, and in 2021 developed some of the basic mathematical foundations and software infrastructure necessary to do so. We immediately found some basic mechanisms in the model that did the kind of things that are essential to interpret language: copying and sequential pattern-matching. We also found some interpretable single neurons, similar to what we found in vision models, which represented various words and concepts. However, we quickly discovered that while _some_ neurons were immediately interpretable, the vast majority were an incoherent pastiche of many different words and concepts. We referred to this phenomenon as _superposition_,7
叠加的基本思想由 Arora 等人在 2016 年描述,更一般地可追溯到关于压缩感知的经典数学工作。叠加假说解释不可解释神经元的想法可追溯到早期关于视觉模型的机械可解释性工作。此时的变化在于,我们清楚地认识到这将成为语言模型的核心问题,比视觉模型严重得多。我们能够提供强有力的理论基础,确信叠加是值得追求的正确假说。
7 The basic idea of superposition was described by Arora _et al_. in 2016, and more generally traces back to classical mathematical work on compressed sensing. The hypothesis that it explained uninterpretable neurons goes back to early mechanistic interpretability work on vision models. What changed at this time was that it became clear this was going to be a central problem for language models, much worse than in vision. We were able to provide a strong theoretical basis for having conviction that superposition was the right hypothesis to pursue.
并且我们很快意识到,模型可能包含数十亿个概念,但以一种混乱混合的方式呈现,我们无法理解。模型使用叠加是因为这允许它表达比神经元数量更多的概念,从而学习更多。如果叠加看起来混乱且难以理解,那是因为,一如既往,AI 模型的学习和运行并未以任何方式优化以对人类可读。
and we quickly realized that the models likely contained billions of concepts, but in a hopelessly mixed-up fashion that we couldn’t make any sense of. The model uses superposition because this allows it to express more concepts than it has neurons, enabling it to learn more. If superposition seems tangled and difficult to understand, that’s because, as ever, the learning and operation of AI models are not optimized in the slightest to be legible to humans.
解释叠加的困难一度阻碍了进展,但最终我们发现(与其他研究者并行)一种来自信号处理的现有技术——稀疏自编码器——可用于找到确实对应更清晰、更人类可理解概念的神经元组合。这些神经元组合能够表达的概念远比单层神经网络中的概念微妙:它们包括“字面或比喻上的回避或犹豫”的概念,以及“表达不满的音乐流派”的概念。我们将这些概念称为特征,并使用稀疏自编码器方法在各种规模的模型中映射它们,包括现代最先进的模型。例如,我们在一个中等规模的商业模型(Claude 3 Sonnet)中发现了超过 3000 万个特征。此外,我们采用了一种称为自动可解释性的方法——即使用 AI 系统本身来分析可解释性特征——以扩展不仅发现特征,而且列出并用人类术语识别其含义的过程。
The difficulty of interpreting superpositions blocked progress for a while, but eventually we discovered (in parallel with others) that an existing technique from signal processing called _sparse autoencoders_ could be used to find _combinations_ of neurons that _did_ correspond to cleaner, more human-understandable concepts. The concepts that these combinations of neurons could express were far more subtle than those of the single-layer neural network: they included the concept of “literally or figuratively hedging or hesitating”, and the concept of “genres of music that express discontent”. We called these concepts _features,_ and used the sparse autoencoder method to map them in models of all sizes, including modern state-of-the-art models. For example, we were able to find over 30 million features in a medium-sized commercial model (Claude 3 Sonnet). Additionally, we employed a method called _autointerpretability_—which uses an AI system itself to analyze interpretability features—to scale the process of not just finding the features, but listing and identifying what they mean in human terms.
发现并识别 3000 万个特征是一个重要的进步,但我们相信即使在一个小型模型中也可能存在十亿或更多的概念,因此我们只发现了可能存在的很小一部分,这方面的工作仍在进行中。更大的模型,例如 Anthropic 最强大产品中使用的模型,则更加复杂。
Finding and identifying 30 million features is a significant step forward, but we believe there may actually be a _billion_ or more concepts in even a small model, so we’ve found only a small fraction of what is probably there, and work in this direction is ongoing. Bigger models, like those used in Anthropic’s most capable products, are more complicated still.
一旦找到特征,我们不仅可以观察其作用,还可以增加或减少其在神经网络处理中的重要性。可解释性的 MRI 可以帮助我们开发和优化干预措施——几乎就像电击大脑的精确部位。最令人难忘的是,我们使用这种方法创建了“金门大桥克劳德”,这是 Anthropic 一个模型的版本,其中“金门大桥”特征被人为放大,导致模型痴迷于这座桥,甚至在无关的对话中也会提起它。
Once a feature is found, we can do more than just observe it in action—we can increase or decrease its importance in the neural network’s processing. The MRI of interpretability can help us develop and refine interventions—almost like zapping a precise part of someone’s brain. Most memorably, we used this method to create “Golden Gate Claude”, a version of one of Anthropic’s models where the “Golden Gate Bridge” feature was artificially amplified, causing the model to become obsessed with the bridge, bringing it up even in unrelated conversations.
最近,我们从追踪和操作特征转向追踪和操作我们称之为“电路”的特征组。这些电路展示了模型思考的步骤:概念如何从输入词中涌现,这些概念如何相互作用形成新概念,以及它们如何在模型内部工作以生成动作。通过电路,我们可以“追踪”模型的思考。例如,如果你问模型“包含达拉斯的州的首府是什么?”,会有一个“位于……内”的电路,导致“达拉斯”特征触发“德克萨斯”特征的激活,然后一个电路导致“奥斯汀”在“德克萨斯”和“首府”之后激活。尽管我们仅通过手动过程发现了少量电路,但我们已经可以用它们来观察模型如何推理问题——例如,在写诗时如何提前规划押韵,以及如何跨语言共享概念。我们正在努力实现电路的自动发现,因为我们预计模型内部有数百万个电路以复杂方式相互作用。
Recently, we’ve moved onward from tracking and manipulating features to tracking and manipulating groups of features that we call “circuits”. These circuits show the steps in a model’s thinking: how concepts emerge from input words, how those concepts interact to form new concepts, and how those work within the model to generate actions. With circuits, we can “trace” the model’s thinking. For example, if you ask the model “What is the capital of the state containing Dallas?”, there is a “located within” circuit that causes the “Dallas” feature to trigger the firing of a “Texas” feature, and then a circuit that causes “Austin” to fire after “Texas” and “capital”. Even though we’ve only found a small number of circuits through a manual process, we can already use them to see how a model reasons through problems—for example how it plans ahead for rhymes when writing poetry, and how it shares concepts across languages. We are working on ways to automate the finding of circuits, as we expect there are millions within a model that interact in complex ways.
所有这些进展,虽然在科学上令人印象深刻,但并未直接回答我们如何利用可解释性来降低我之前列出的风险。假设我们已经识别出一堆概念和电路——甚至假设我们知道所有概念和电路,并且能比今天更好地理解和组织它们。那又怎样?我们如何利用这一切?从抽象理论到实际价值之间仍然存在差距。
All of this progress, while scientifically impressive, doesn’t directly answer the question of how we can use interpretability to reduce the risks I listed earlier. Suppose we have identified a bunch of concepts and circuits—suppose, even, that we know all of them, and we can understand and organize them much better than we can today. So what? How do we _use_ all of it? There’s still a gap from abstract theory to practical value.
为了帮助缩小这一差距,我们已开始尝试使用我们的可解释性方法来发现和诊断模型中的问题。最近,我们做了一个实验,让一个“红队”故意在模型中引入对齐问题(例如,模型倾向于利用任务中的漏洞),并让多个“蓝队”负责找出模型的问题所在。多个蓝队成功了;特别相关的是,其中一些蓝队在调查过程中有效地应用了可解释性工具。我们仍需扩展这些方法,但这次练习帮助我们在使用可解释性技术发现和解决模型缺陷方面获得了实际经验。
To help close that gap, we’ve begun experimenting with using our interpretability methods to find and diagnose problems in models. Recently, we did an experiment where we had a “red team” deliberately introduce an alignment issue into a model (say, a tendency for the model to exploit a loophole in a task) and gave various “blue teams” the task of figuring out what was wrong with it. Multiple blue teams succeeded; of particular relevance here, some of them productively applied interpretability tools during the investigation. We still need to scale these methods, but the exercise helped us gain some practical experience using interpretability techniques to find and address flaws in our models.
我们的长期目标是能够观察最先进的模型,并基本上进行“脑部扫描”:一种检查,有很高的概率识别出各种问题,包括撒谎或欺骗的倾向、追求权力、越狱漏洞、模型整体的认知优势和劣势等等。这将与各种训练和对齐模型的技术结合使用,有点像医生可能进行核磁共振成像(MRI)来诊断疾病,然后开药治疗,再进行另一次 MRI 来查看治疗进展,如此循环。
Our long-run aspiration is to be able to look at a state-of-the-art model and essentially do a “brain scan”: a checkup that has a high probability of identifying a wide range of issues including tendencies to lie or deceive, power-seeking, flaws in jailbreaks, cognitive strengths and weaknesses of the model as a whole, and much more. This would then be used in tandem with the various techniques for training and aligning models, a bit like how a doctor might do an MRI to diagnose a disease, then prescribe a drug to treat it, then do another MRI to see how the treatment is progressing, and so on8
一种说法是,可解释性应像模型对齐的_测试集_一样发挥作用,而传统的对齐技术,如可扩展监督、基于人类反馈的强化学习(RLHF)、宪法 AI 等,则应像_训练集_一样发挥作用。也就是说,可解释性作为模型对齐的独立检查,不受训练过程的影响,训练过程可能激励模型_看起来_对齐而实际上并非如此。这种观点的两个后果是:(a)在生产中,我们应非常谨慎地直接训练或优化可解释性输出(特征/概念、电路),因为这会破坏其信号的独立性;(b)重要的是,在单次生产运行中,不要“使用”诊断测试信号_太多次_来告知训练过程的变更,因为这会将关于独立测试信号的信息逐渐泄露给训练过程(尽管比(a)慢得多)。换句话说,我们建议在评估正式的、高风险的生产模型时,对待可解释性分析要像对待隐藏评估或测试集一样谨慎。
8 One way to say this is that interpretability should function like the _test set_ for model alignment, while traditional alignment techniques such as scalable supervision, RLHF, constitutional AI, etc. should function as the _training set_. That is, interpretability acts as an independent check on the alignment of models, uncontaminated by the training process which might incentivize models to _appear_ aligned without being so. Two consequences of this view are that (a) we should be very hesitant to directly train or optimize on interpretability outputs (features/concepts, circuits) in production, as this destroys the independence of their signal, and (b) it’s important not to “use” the diagnostic test signal _too many times_ in one production run to inform changes to the training process, as this gradually leaks bits of information about the independent test signal to the training process (though much more slowly than (a)). In other words, we recommend that in assessing official, high-stakes production models, we treat interpretability analysis with the same care we would treat a hidden evaluation or test set.
很可能,我们测试和部署最强大模型(例如,在我们的负责任扩展政策框架中处于 AI 安全级别 4 的模型)的关键部分,将是执行和形式化此类测试。
. It is likely that a key part of how we will test and deploy the most capable models (for example, those at AI Safety Level 4 in our Responsible Scaling Policy framework) is by performing and formalizing such tests.
一方面,最近的进展——尤其是关于电路和基于可解释性的模型测试的结果——让我觉得我们正处于以重大方式攻克可解释性的边缘。尽管摆在我们面前的任务艰巨,但我能看到一条现实的道路,使可解释性成为一种复杂而可靠的方法,用于诊断即使是极其先进的 AI 中的问题——真正的“AI 的 MRI”。事实上,按照目前的轨迹,我强烈押注可解释性将在 5-10 年内达到这一点。
On one hand, recent progress—especially the results on circuits and on interpretability-based testing of models—has made me feel that we are on the verge of cracking interpretability in a big way. Although the task ahead of us is Herculean, I can see a realistic path towards interpretability being a sophisticated and reliable way to diagnose problems in even very advanced AI—a true “MRI for AI”. In fact, on its current trajectory I would bet strongly in favor of interpretability reaching this point within 5-10 years.
另一方面,我担心 AI 本身发展如此之快,以至于我们可能连这么多时间都没有。正如我在其他地方所写,我们可能在 2026 或 2027 年就拥有相当于“数据中心里的天才之国”的 AI 系统。我非常担心在未能更好地掌握可解释性的情况下部署这样的系统。这些系统将绝对成为经济、技术和国家安全的核心,并且将具备如此大的自主性,以至于我认为人类完全不了解它们的工作原理基本上是难以接受的。
On the other hand, I worry that AI itself is advancing so quickly that we might not have even this much time. As I’ve written elsewhere, we could have AI systems equivalent to a “country of geniuses in a datacenter” as soon as 2026 or 2027. I am very concerned about deploying such systems without a better handle on interpretability. These systems will be absolutely central to the economy, technology, and national security, and will be capable of so much autonomy that I consider it basically unacceptable for humanity to be totally ignorant of how they work.
因此,我们正处于可解释性与模型智能之间的竞赛。这不是一个全有或全无的问题:正如我们所看到的,可解释性的每一次进步都定量地提高了我们观察模型内部并诊断其问题的能力。我们取得的进步越多,“数据中心里的天才之国”顺利运行的可能性就越大。AI 公司、研究人员、政府和社会可以做几件事来改变局面:
We are thus in a race between interpretability and model intelligence. It is not an all-or-nothing matter: as we’ve seen, every advance in interpretability quantitatively increases our ability to look inside models and diagnose their problems. The more such advances we have, the greater the likelihood that the “country of geniuses in a datacenter” goes well. There are several things that AI companies, researchers, governments, and society can do to tip the scales:
首先,公司、学术界或非营利组织的 AI 研究人员可以直接从事可解释性研究来加速其发展。可解释性得到的关注远少于不断涌现的模型发布,但它可以说更为重要。对我来说,这也是加入该领域的理想时机:最近的“电路”结果同时开辟了许多方向。Anthropic 正在加倍投入可解释性,我们的目标是到 2027 年实现“可解释性能够可靠地检测大多数模型问题”。我们也在投资可解释性初创公司。
First, AI researchers in companies, academia, or nonprofits can accelerate interpretability by directly working on it. Interpretability gets less attention than the constant deluge of model releases, but it is arguably more important. It also feels to me like it is an ideal time to join the field: the recent “circuits” results have opened up many directions in parallel. Anthropic is doubling down on interpretability, and we have a goal of getting to “interpretability can reliably detect most model problems” by 2027. We are also investing in interpretability startups.
但如果这是整个科学界的共同努力,成功的几率会更大。其他公司,如 Google DeepMind 和 OpenAI,也有一些可解释性的努力,但我强烈鼓励它们分配更多资源。如果有所帮助,Anthropic 将尝试将可解释性商业化,以创造独特优势,尤其是在那些决策解释能力至关重要的行业。如果你是竞争对手,不希望这种情况发生,那么你也应该加大对可解释性的投入!
But the chances of succeeding at this are greater if it is an effort that spans the whole scientific community. Other companies, such as Google DeepMind and OpenAI, have some interpretability efforts, but I strongly encourage them to allocate more resources. If it helps, Anthropic will be trying to apply interpretability commercially to create a unique advantage, especially in industries where the ability to provide an explanation for decisions is at a premium. If you are a competitor and you don’t want this to happen, you too should invest more in interpretability!
可解释性也自然适合学术和独立研究人员:它具有基础科学的味道,并且许多部分可以在不需要巨大计算资源的情况下进行研究。需要明确的是,一些独立研究人员和学者确实在研究可解释性,但我们需要更多。
Interpretability is also a natural fit for academic and independent researchers: it has the flavor of basic science, and many parts of it can be studied without needing huge computational resources. To be clear, some independent researchers and academics do work on interpretability, but we need many more9
其次,政府可以使用轻触式规则来鼓励可解释性研究的发展及其在解决前沿 AI 模型问题中的应用。鉴于“AI MRI”实践尚处于萌芽和未开发阶段,为什么现阶段监管或强制要求公司进行这些实践没有意义,原因应该很清楚:甚至不清楚一项潜在法律应该要求公司做什么。但要求公司透明地披露其安全实践(其负责任扩展政策及其执行情况),包括如何在发布前使用可解释性测试模型,将允许公司相互学习,同时明确谁的行为更负责任,促进“逐顶竞赛”。我们在给加州前沿模型工作组的回应中建议将安全/负责任扩展政策透明度作为加州法律的可能方向(工作组本身也提到了类似的想法)。这一概念也可以推广到联邦层面或其他国家。
9 Bizarrely, mechanistic interpretability sometimes seems to meet substantial cultural resistance in academia. For example, I am concerned by reports that a very popular mechanistic interpretability ICML conference workshop was rejected on seemingly pretextual grounds. If true, this behavior is shortsighted and self-defeating at exactly a time when academics in AI are looking for ways to maintain relevance.
第三,政府可以使用出口管制来创建一个“安全缓冲”,这可能为可解释性争取更多时间,以便在我们达到最强大的 AI 之前取得进展。我一直支持对华芯片出口管制,因为我相信民主国家必须在 AI 方面保持对专制国家的领先。但这些政策还有一个额外的好处。如果美国和其他民主国家在接近“数据中心里的天才之国”时拥有明显的领先优势,我们或许可以“花费”部分领先优势来确保可解释性在进入真正强大的 AI 之前有更坚实的基础,同时仍然击败我们的专制对手。即使是一两年的领先优势——我相信有效且执行良好的出口管制可以给我们带来——可能意味着在达到变革性能力水平时,“AI MRI”基本有效与无效之间的区别。一年前,我们还无法追踪神经网络的思维,也无法识别其中的数百万个概念;今天我们可以。相比之下,如果美国和中国同时达到强大的 AI(这是我预计在没有出口管制的情况下会发生的情况),地缘政治激励将使任何放缓几乎不可能。
. Finally, if you are in another scientific field and are looking for new opportunities, interpretability may be a promising bet, as it offers rich data, exciting burgeoning methods, and enormous real-world value. Neuroscientists especially should consider this, as it’s much easier to collect data on artificial neural networks than biological ones, and some of the conclusions can be applied back to neuroscience. If you're interested in joining Anthropic's Interpretability team, we have open Research Scientist and Research Engineer roles.
所有这些——加速可解释性、轻触式透明度立法以及对华芯片出口管制——本身都是好主意,几乎没有重大缺点。我们无论如何都应该做所有这些事情。但当我们意识到它们可能决定可解释性是在强大 AI 之前还是之后被解决时,它们就变得更加重要。
Second, governments can uselight-touch rulesto encourage the development of interpretability research and its application to addressing problems with frontier AI models. Given how nascent and undeveloped the practice of “AI MRI” is, it should be clear why it doesn’t make sense to regulate or mandate that companies conduct them, at least at this stage: it’s not even clear _what_ a prospective law should ask companies to do. But a requirement for companies to transparently disclose their safety and security practices (their Responsible Scaling Policy, or RSP, and its execution), including how they’re using interpretability to test models before release, would allow companies to learn from each other while also making clear who is behaving more responsibly, fostering a “race to the top”. We’ve suggested safety/security/RSP transparency as a possible direction for California law in our response to the California frontier model task force (which itself mentions some of the same ideas). This concept could also be exported federally, or to other countries.
强大的 AI 将塑造人类的命运,我们理应在它们彻底改变我们的经济、生活和未来之前理解我们自己的创造。
Third, governments can use export controls to create a “security buffer” that might give interpretability more time to advance before we reach the most powerful AI. I’ve long been a proponent of export controls on chips to China because I believe that democratic countries must remain ahead of autocracies in AI. But these policies also have an additional benefit. If the US and other democracies have a clear lead in AI as they approach the “country of geniuses in a datacenter”, we may be able to “spend” a portion of that lead to ensure interpretability10
感谢 Tom McGrath、Martin Wattenberg、Chris Olah、Ben Buchanan 以及 Anthropic 内部的许多人对本文草稿的反馈。
10 Along with other techniques for mitigating risk, of course—I don’t intend to imply that interpretability is our _only_ risk mitigation tool.
is on a more solid footing before proceeding to truly powerful AI, while still defeating our authoritarian adversaries11
11 I am in fact quite skeptical that any slowdown to address risk is possible even among companies within democratic countries, given the incredible economic value of AI. Fighting the market head-on like this feels like trying to stop a freight train with your toe. But if truly compelling evidence of the dangers of autonomous AI emerged, I think it would be just barely possible. Contrary to the claims of advocates, I don’t think truly compelling evidence exists today, and I actually think the most likely route for providing “smoking gun” evidence of danger is interpretability itself—yet another reason to invest in it!
. Even a 1- or 2-year lead, which I believe effective and well-enforced export controls can give us, could mean the difference between an “AI MRI” that essentially works when we reach transformative capability levels, and one that does not. One year ago we couldn’t trace the thoughts of a neural network and couldn’t identify millions of concepts inside them; today we can. By contrast, if the US and China reach powerful AI simultaneously (which is what I expect to happen without export controls), the geopolitical incentives will make any slowdown at all essentially impossible.
All of these—accelerating interpretability, light-touch transparency legislation, and export controls on chips to China—have the virtue of being good ideas in their own right, with few meaningful downsides. We should do all of them anyway. But they become even more important when we realize that they might make the difference between interpretability being solved before powerful AI or after it.
Powerful AI will shape humanity’s destiny, and we deserve to understand our own creations _before_ they radically transform our economy, our lives, and our future.
_Thanks to Tom McGrath, Martin Wattenberg, Chris Olah, Ben Buchanan, and many people within Anthropic for feedback on drafts of this article._