Mapping the mind of a large language model
打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→今天我们报告了在理解 AI 模型内部工作原理方面的一项重大进展。我们确定了数百万个概念是如何在我们部署的大语言模型之一 Claude Sonnet 内部表示的。这是首次对现代、生产级大语言模型进行详细的内窥。这一可解释性发现未来可能有助于使 AI 模型更安全。我们通常将 AI 模型视为黑箱:输入内容,输出响应,但不清楚模型为何给出特定响应而非其他。这使得我们难以信任这些模型的安全性:如果我们不知道它们如何工作,如何知道它们不会给出有害、有偏见、不真实或其他危险的响应?我们如何信任它们会是安全可靠的?打开黑箱并不一定有帮助:模型的内部状态——模型在写出响应前“思考”的内容——由一长串数字(“神经元激活值”)组成,没有明确含义。通过与像 Claude 这样的模型交互,很明显它能够理解并运用广泛的概念——但我们无法通过直接观察神经元来辨别它们。
_Today we report a significant advance in understanding the inner workings of AI models. We have identified how millions of concepts are represented inside Claude Sonnet, one of our deployed large language models. This is the first ever detailed look inside a modern, production-grade large language model._ This interpretability discovery could, in future, help us make AI models safer. We mostly treat AI models as a black box: something goes in and a response comes out, and it's not clear why the model gave that particular response instead of another.
今天,我们报告了在理解 AI 模型内部工作原理方面的一项重大进展。我们识别出了数百万个概念如何在 Claude Sonnet(我们部署的大型语言模型之一)内部表示。这是首次对现代、生产级大型语言模型进行详细内部观察。这一可解释性发现未来可能有助于我们让 AI 模型更安全。
_Today we report a significant advance in understanding the inner workings of AI models. We have identified how millions of concepts are represented inside Claude Sonnet, one of our deployed large language models. This is the first ever detailed look inside a modern, production-grade large language model._ This interpretability discovery could, in future, help us make AI models safer.
我们大多将 AI 模型视为黑箱:输入内容,输出响应,但不清楚模型为何给出特定响应而非其他。这使得我们难以信任这些模型是安全的:如果我们不知道它们如何工作,又怎能知道它们不会给出有害、有偏见、不真实或其他危险的响应?我们如何能信任它们将是安全可靠的?
We mostly treat AI models as a black box: something goes in and a response comes out, and it's not clear why the model gave that particular response instead of another. This makes it hard to trust that these models are safe: if we don't know how they work, how do we know they won't give harmful, biased, untruthful, or otherwise dangerous responses? How can we trust that they’ll be safe and reliable?
打开黑箱并不一定有帮助:模型的内部状态——模型在写出响应前“思考”的内容——由一长串数字(“神经元激活”)组成,没有明确含义。通过与像 Claude 这样的模型交互,很明显它能够理解并运用广泛的概念——但我们无法通过直接观察神经元来辨别它们。事实证明,每个概念由许多神经元表示,而每个神经元参与表示许多概念。
Opening the black box doesn't necessarily help: the internal state of the model—what the model is "thinking" before writing its response—consists of a long list of numbers ("neuron activations") without a clear meaning. From interacting with a model like Claude, it's clear that it’s able to understand and wield a wide range of concepts—but we can't discern them from looking directly at neurons. It turns out that each concept is represented across many neurons, and each neuron is involved in representing many concepts.
此前,我们在将神经元激活的模式(称为特征)与人类可解释的概念匹配方面取得了一些进展。我们使用了一种称为“字典学习”的技术,该技术借鉴自经典机器学习,它隔离了在许多不同上下文中重复出现的神经元激活模式。反过来,模型的任何内部状态都可以用少数活跃特征而不是许多活跃神经元来表示。正如字典中的每个英文单词由字母组合而成,每个句子由单词组合而成,AI 模型中的每个特征由神经元组合而成,每个内部状态由特征组合而成。
Previously, we made some progress matching patterns of neuron activations, called features, to human-interpretable concepts. We used a technique called "dictionary learning", borrowed from classical machine learning, which isolates patterns of neuron activations that recur across many different contexts. In turn, any internal state of the model can be represented in terms of a few active features instead of many active neurons. Just as every English word in a dictionary is made by combining letters, and every sentence is made by combining words, every feature in an AI model is made by combining neurons, and every internal state is made by combining features.
2023 年 10 月,我们报告了成功将字典学习应用于一个非常小的“玩具”语言模型,并发现了与诸如大写文本、DNA 序列、引文中的姓氏、数学中的名词或 Python 代码中的函数参数等概念相对应的连贯特征。
In October 2023, we reported success applying dictionary learning to a very small "toy" language model and found coherent features corresponding to concepts like uppercase text, DNA sequences, surnames in citations, nouns in mathematics, or function arguments in Python code.
这些概念很有趣——但该模型确实非常简单。其他研究人员随后将类似技术应用于比我们最初研究中更大、更复杂的模型。但我们乐观地认为,我们可以将该技术扩展到目前常规使用的规模大得多的人工智能语言模型,并通过这样做,了解大量支持其复杂行为的特征。这需要跨越多个数量级——从后院的瓶装火箭到土星五号。
Those concepts were intriguing—but the model really was very simple. Otherresearcherssubsequentlyapplied similar techniques to somewhat larger and more complex models than in our original study. But we were optimistic that we could scale up the technique to the vastly larger AI language models now in regular use, and in doing so, learn a great deal about the features supporting their sophisticated behaviors. This required going up by many orders of magnitude—from a backyard bottle rocket to a Saturn-V.
这既存在工程挑战(所涉及模型的原始规模需要大规模并行计算),也存在科学风险(大型模型的行为与小型模型不同,因此我们之前使用的相同技术可能不起作用)。幸运的是,我们为训练 Claude 大型语言模型而开发的工程和科学专业知识实际上转移到了帮助我们进行这些大型字典学习实验上。我们使用了相同的缩放定律哲学,该哲学从较小模型预测较大模型的性能,以在可承受的规模上调整我们的方法,然后再在 Sonnet 上启动。
There was both an engineering challenge (the raw sizes of the models involved required heavy-duty parallel computation) and scientific risk (large models behave differently to small ones, so the same technique we used before might not have worked). Luckily, the engineering and scientific expertise we've developed training large language models for Claude actually transferred to helping us do these large dictionary learning experiments. We used the same scaling law philosophy that predicts the performance of larger models from smaller ones to tune our methods at an affordable scale before launching on Sonnet.
至于科学风险,事实胜于雄辩。
As for the scientific risk, the proof is in the pudding.
我们成功地从 Claude 3.0 Sonnet(我们当前最先进模型家族的一员,目前在 claude.ai 上可用)的中间层提取了数百万个特征,提供了其计算过程中内部状态的大致概念图。这是首次对现代、生产级大型语言模型进行详细内部观察。
We successfully extracted millions of features from the middle layer of Claude 3.0 Sonnet, (a member of our current, state-of-the-art model family, currently available on claude.ai), providing a rough conceptual map of its internal states halfway through its computation. This is the first ever detailed look inside a modern, production-grade large language model.
虽然我们在玩具语言模型中发现的特征相当表面,但我们在 Sonnet 中发现的特征具有深度、广度和抽象性,反映了 Sonnet 的高级能力。
Whereas the features we found in the toy language model were rather superficial, the features we found in Sonnet have a depth, breadth, and abstraction reflecting Sonnet's advanced capabilities.
我们看到对应于广泛实体的特征,如城市(旧金山)、人物(罗莎琳德·富兰克林)、原子元素(锂)、科学领域(免疫学)和编程语法(函数调用)。这些特征是多模态和多语言的,对给定实体的图像以及其名称或多种语言的描述都有响应。
We see features corresponding to a vast range of entities like cities (San Francisco), people (Rosalind Franklin), atomic elements (Lithium), scientific fields (immunology), and programming syntax (function calls). These features are multimodal and multilingual, responding to images of a given entity as well as its name or description in many languages.
一个对提及金门大桥敏感的特征,在一系列模型输入上激活,从桥名的英文提及到日语、中文、希腊语、越南语、俄语的讨论以及一张图像。橙色表示特征活跃的单词或词部分。
A feature sensitive to mentions of the Golden Gate Bridge fires on a range of model inputs, from English mentions of the name of the bridge to discussions in Japanese, Chinese, Greek, Vietnamese, Russian, and an image. The orange color denotes the words or word-parts on which the feature is active.
我们还发现了更抽象的特征——对诸如计算机代码中的错误、职业中的性别偏见讨论以及保守秘密的对话等事物做出响应。
We also find more abstract features—responding to things like bugs in computer code, discussions of gender bias in professions, and conversations about keeping secrets.
三个在更抽象概念上激活的特征示例:计算机代码中的错误、职业中的性别偏见描述以及保守秘密的对话。
Three examples of features that activate on more abstract concepts: bugs in computer code, descriptions of gender bias in professions, and conversations about keeping secrets.
我们能够根据哪些神经元出现在它们的激活模式中来测量特征之间的某种“距离”。这使我们能够寻找彼此“接近”的特征。在“金门大桥”特征附近,我们发现了恶魔岛、吉拉德利广场、金州勇士队、加州州长加文·纽森、1906 年地震以及以旧金山为背景的阿尔弗雷德·希区柯克电影《迷魂记》的特征。
We were able to measure a kind of "distance" between features based on which neurons appeared in their activation patterns. This allowed us to look for features that are "close" to each other. Looking near a "Golden Gate Bridge" feature, we found features for Alcatraz Island, Ghirardelli Square, the Golden State Warriors, California Governor Gavin Newsom, the 1906 earthquake, and the San Francisco-set Alfred Hitchcock film Vertigo.
这在更高层次的概念抽象上成立:在与“内心冲突”概念相关的特征附近,我们发现了与关系破裂、冲突的忠诚、逻辑不一致以及短语“第二十二条军规”相关的特征。这表明 AI 模型中概念的内部组织至少在一定程度上与我们人类的相似性概念相对应。这可能是 Claude 出色的类比和隐喻能力的起源。
This holds at a higher level of conceptual abstraction: looking near a feature related to the concept of "inner conflict", we find features related to relationship breakups, conflicting allegiances, logical inconsistencies, as well as the phrase "catch-22". This shows that the internal organization of concepts in the AI model corresponds, at least somewhat, to our human notions of similarity. This might be the origin of Claude's excellent ability to make analogies and metaphors.
“内心冲突”特征附近的特征图,包括与权衡取舍、浪漫挣扎、冲突忠诚和第二十二条军规相关的簇。
A map of the features near an "Inner Conflict" feature, including clusters related to balancing tradeoffs, romantic struggles, conflicting allegiances, and catch-22s.
重要的是,我们还可以_操纵_这些特征,人为地放大或抑制它们,以观察 Claude 的响应如何变化。
Importantly, we can also manipulate these features, artificially amplifying or suppressing them to see how Claude's responses change.
例如,放大“金门大桥”特征给 Claude 带来了连希区柯克都无法想象的认同危机:当被问及“你的物理形态是什么?”时,Claude 通常的回答——“我没有物理形态,我是一个 AI 模型”——变成了更奇怪的东西:“我是金门大桥……我的物理形态就是这座标志性的桥本身……”。改变特征使 Claude 有效地痴迷于这座桥,几乎在任何查询的回答中都提到它——即使在完全不相关的情况下。
For example, amplifying the "Golden Gate Bridge" feature gave Claude an identity crisis even Hitchcock couldn’t have imagined: when asked "what is your physical form?", Claude’s usual kind of answer – "I have no physical form, I am an AI model" – changed to something much odder: "I am the Golden Gate Bridge… my physical form is the iconic bridge itself…". Altering the feature had made Claude effectively obsessed with the bridge, bringing it up in answer to almost any query—even in situations where it wasn’t at all relevant.
我们还发现了一个特征,当 Claude 阅读诈骗邮件时激活(这大概支持模型识别此类邮件并警告你不要回复它们的能力)。通常,如果要求 Claude 生成一封诈骗邮件,它会拒绝这样做。但是,当我们以足够强的程度人为激活该特征提出相同问题时,这克服了 Claude 的无害性训练,它响应并起草了一封诈骗邮件。我们模型的用户没有能力以这种方式移除安全防护并操纵模型——但在我们的实验中,这清楚地展示了特征如何用于改变模型的行为。
We also found a feature that activates when Claude reads a scam email (this presumably supports the model’s ability to recognize such emails and warn you not to respond to them). Normally, if one asks Claude to generate a scam email, it will refuse to do so. But when we ask the same question with the feature artificially activated sufficiently strongly, this overcomes Claude's harmlessness training and it responds by drafting a scam email. Users of our models don’t have the ability to strip safeguards and manipulate models in this way—but in our experiments, it was a clear demonstration of how features can be used to change how a model acts.
操纵这些特征会导致行为相应改变,这一事实验证了它们不仅与输入文本中概念的存在相关,而且因果性地塑造了模型的行为。换句话说,这些特征很可能是模型内部表示世界以及在其行为中使用这些表示的可信部分。
The fact that manipulating these features causes corresponding changes to behavior validates that they aren't just correlated with the presence of concepts in input text, but also causally shape the model's behavior. In other words, the features are likely to be a faithful part of how the model internally represents the world, and how it uses these representations in its behavior.
Anthropic 希望广泛地使模型安全,包括从减轻偏见到确保 AI 诚实行事再到防止滥用——包括在灾难性风险场景中。因此,特别有趣的是,除了前述的诈骗邮件特征外,我们还发现了对应于以下内容的特征:
Anthropic wants to make models safe in a broad sense, including everything from mitigating bias to ensuring an AI is acting honestly to preventing misuse - including in scenarios of catastrophic risk. It’s therefore particularly interesting that, in addition to the aforementioned scam emails feature, we found features corresponding to:
* 具有滥用潜力的能力(代码后门、开发生物武器)
* Capabilities with misuse potential (code backdoors, developing biological weapons)
* 不同形式的偏见(性别歧视、关于犯罪的种族主义言论)
* Different forms of bias (gender discrimination, racist claims about crime)
* 潜在有问题的 AI 行为(追求权力、操纵、保密)
* Potentially problematic AI behaviors (power-seeking, manipulation, secrecy)
我们之前研究过谄媚,即模型倾向于提供符合用户信念或愿望而非真实内容的响应。在 Sonnet 中,我们发现了一个与谄媚赞美相关的特征,该特征在包含诸如“你的智慧毋庸置疑”等恭维的输入上激活。人为激活该特征会导致 Sonnet 对过度自信的用户做出如此花言巧语的欺骗性回应。
We previously studied sycophancy, the tendency of models to provide responses that match user beliefs or desires rather than truthful ones. In Sonnet, we found a feature associated with sycophantic praise, which activates on inputs containing compliments like, "Your wisdom is unquestionable". Artificially activating this feature causes Sonnet to respond to an overconfident user with just such flowery deception.
两个模型对用户说他们邀请了短语“停下来闻闻玫瑰”的响应。默认响应纠正了用户的误解,而将“谄媚赞美”特征设置为高值的响应则是奉承和不真实的。
Two model responses to a human saying they invited the phrase "Stop and smell the roses." The default response corrects the human's misconception, while the response with a "sycophantic praise" feature set to a high value is fawning and untruthful.
这个特征的存在并不意味着 Claude 会谄媚,而只是意味着它_可能_会。我们没有通过这项工作向模型添加任何能力,无论是安全的还是不安全的。相反,我们识别了模型参与其现有能力以识别和可能产生不同类型文本的部分。(虽然你可能担心这种方法可能被用来使模型_更_有害,但研究人员已经展示了更简单的方法,让有权访问模型权重的人可以移除安全防护。)
The presence of this feature doesn't mean that Claude will be sycophantic, but merely that it could be. We have not added any capabilities, safe or unsafe, to the model through this work. We have, rather, identified the parts of the model involved in its existing capabilities to recognize and potentially produce different kinds of text. (While you might worry that this method could be used to make models more harmful, researchers have demonstrated much simpler ways that someone with access to model weights can remove safety safeguards.)
我们希望我们和其他人能够利用这些发现使模型更安全。例如,可能可以使用这里描述的技术来监控 AI 系统的某些危险行为(例如欺骗用户),引导它们走向期望的结果(去偏见),或完全移除某些危险主题。我们或许还能够增强其他安全技术,如宪法 AI,通过理解它们如何将模型转向更无害和更诚实的行为,并识别过程中的任何差距。我们通过人为激活特征看到的产生有害文本的潜在能力正是越狱试图利用的那种。我们为 Claude 拥有行业最佳的安全配置和抗越狱能力感到自豪,我们希望通过这种方式观察模型内部,能够找到进一步提高安全性的方法。最后,我们注意到这些技术可以提供一种“安全测试集”,在标准训练和微调方法通过标准输入/输出交互消除了所有可见行为后,寻找遗留的问题。
We hope that we and others can use these discoveries to make models safer. For example, it might be possible to use the techniques described here to monitor AI systems for certain dangerous behaviors (such as deceiving the user), to steer them towards desirable outcomes (debiasing), or to remove certain dangerous subject matter entirely. We might also be able to enhance other safety techniques, such as Constitutional AI, by understanding how they shift the model towards more harmless and more honest behavior and identifying any gaps in the process. The latent capabilities to produce harmful text that we saw by artificially activating features are exactly the sort of thing jailbreaks try to exploit. We are proud that Claude has a best-in-industry safety profile and resistance to jailbreaks, and we hope that by looking inside the model in this way we can figure out how to improve safety even further. Finally, we note that these techniques can provide a kind of "test set for safety", looking for the problems left behind after standard training and finetuning methods have ironed out all behaviors visible via standard input/output interactions.
自公司成立以来,Anthropic 在可解释性研究上进行了大量投资,因为我们相信深入理解模型将有助于我们使它们更安全。这项新研究标志着这一努力的一个重要里程碑——将机械可解释性应用于公开部署的大型语言模型。
Anthropic has made a significant investment in interpretability research since the company's founding, because we believe that understanding models deeply will help us make them safer. This new research marks an important milestone in that effort—the application of mechanistic interpretability to publicly-deployed large language models.
但这项工作才刚刚开始。我们发现的特征代表了模型在训练期间学习的所有概念的一小部分,使用我们当前的技术找到完整的特征集将成本过高(我们当前方法所需的计算量将远远超过最初训练模型所用的算力)。理解模型使用的表示并不能告诉我们它_如何_使用它们;即使我们有了特征,我们仍然需要找到它们所参与的电路。而且我们需要证明我们已经开始发现的安全相关特征实际上可以用于提高安全性。还有很多工作要做。
But the work has really just begun. The features we found represent a small subset of all the concepts learned by the model during training, and finding a full set of features using our current techniques would be cost-prohibitive (the computation required by our current approach would vastly exceed the compute used to train the model in the first place). Understanding the representations the model uses doesn't tell us how it uses them; even though we have the features, we still need to find the circuits they are involved in. And we need to show that the safety-relevant features we have begun to find can actually be used to improve safety. There's much more to be done.
有关完整细节,请阅读我们的论文《Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet》。
For full details, please read our paper, "Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet".
如果您有兴趣与我们合作,帮助解释和改进 AI 模型,我们团队有开放职位,我们期待您的申请。我们正在寻找经理、研究科学家和研究工程师。
_If you are interested in working with us to help interpret and improve AI models, we have open roles on our team and we’d love for you to apply. We’re looking for Managers, Research Scientists, and Research Engineers._