追踪大型语言模型的思维

Tracing the thoughts of a large language model

Anthropic Anthropic · Anthropic · 2025-03-27 · Anthropic Research ↗

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

像 Claude 这样的语言模型并非由人类直接编程,而是通过大量数据进行训练。在训练过程中,它们学会了解决问题的策略。这些策略被编码在模型为每个词所执行的数十亿次计算中,对我们(模型的开发者)来说难以理解。这意味着我们并不了解模型完成大多数任务的方式。了解 Claude 这样的模型如何“思考”,将使我们能够更好地理解它们的能力,并帮助确保它们按照我们的意图行事。例如:Claude 能说几十种语言。它在“头脑中”使用的是什么语言(如果有的话)?

Language models like Claude aren't programmed directly by humans—instead, they‘re trained on large amounts of data. During that training process, they learn their own strategies to solve problems. These strategies are encoded in the billions of computations a model performs for every word it writes. They arrive inscrutable to us, the model’s developers. This means that we don’t understand how models do most of the things they do. Knowing how models like Claude think would allow us to have a better understanding of their abilities, as well as help us ensure that they’re doing what we intend them to. For example: * Claude can speak dozens of languages.

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 9)

全文 · Full text(逐段中英对照)

概述 Overview

像 Claude 这样的语言模型并非由人类直接编程——相反,它们是通过大量数据进行训练的。在训练过程中,它们学会了解决问题的策略。这些策略编码在模型为每个词所执行的数十亿次计算中。对我们这些模型开发者来说,这些计算是难以理解的。这意味着我们并不了解模型所做的大部分事情是如何完成的。

Language models like Claude aren't programmed directly by humans—instead, they‘re trained on large amounts of data. During that training process, they learn their own strategies to solve problems. These strategies are encoded in the billions of computations a model performs for every word it writes. They arrive inscrutable to us, the model’s developers. This means that we don’t understand how models do most of the things they do.

了解 Claude 这样的模型是如何“思考”的,将使我们能够更好地理解它们的能力,并帮助我们确保它们按照我们的意图行事。例如:

Knowing how models like Claude think would allow us to have a better understanding of their abilities, as well as help us ensure that they’re doing what we intend them to. For example:

* Claude 能说几十种语言。它在“头脑中”使用的是哪种语言(如果有的话)?

* Claude can speak dozens of languages. What language, if any, is it using "in its head"?

* Claude 一次写一个词。它只专注于预测下一个词,还是会提前规划?

* Claude writes text one word at a time. Is it only focusing on predicting the next word or does it ever plan ahead?

* Claude 可以逐步写出它的推理过程。这种解释是否代表了它得出答案的实际步骤,还是有时它只是在为一个已成定论的结论编造一个看似合理的论证?

* Claude can write out its reasoning step-by-step. Does this explanation represent the actual steps it took to get to an answer, or is it sometimes fabricating a plausible argument for a foregone conclusion?

我们从神经科学领域汲取灵感,该领域长期研究思考生物体的复杂内部机制,并试图构建一种 AI 显微镜,让我们能够识别活动模式和信息流。仅仅通过与 AI 模型对话所能学到的东西是有限的——毕竟,人类(甚至神经科学家)也不完全了解我们自己大脑运作的所有细节。因此,我们深入内部观察。

We take inspiration from the field of neuroscience, which has long studied the messy insides of thinking organisms, and try to build a kind of AI microscope that will let us identify patterns of activity and flows of information. There are limits to what you can learn just by talking to an AI model—after all, humans (even neuroscientists) don't know all the details of how our own brains work. So we look inside.

今天,我们分享两篇新论文,它们代表了“显微镜”开发的进展,以及将其应用于观察新的“AI 生物学”。在第一篇论文中,我们将先前在模型内部定位可解释概念(“特征”)的工作扩展到将这些概念连接成计算“回路”,揭示了将输入 Claude 的词语转换为输出词语的部分路径。在第二篇论文中,我们深入研究了 Claude 3.5 Haiku,对代表十个关键模型行为的简单任务进行了深入研究,包括上述三个任务。我们的方法揭示了 Claude 响应这些提示时发生的一部分情况,这足以看到确凿的证据表明:

Today, we're sharing two new papers that represent progress on the development of the "microscope", and the application of it to see new "AI biology". In the first paper, we extend our prior work locating interpretable concepts ("features") inside a model to link those concepts together into computational "circuits", revealing parts of the pathway that transforms the words that go into Claude into the words that come out. In the second, we look inside Claude 3.5 Haiku, performing deep studies of simple tasks representative of ten crucial model behaviors, including the three described above. Our method sheds light on a part of what happens when Claude responds to these prompts, which is enough to see solid evidence that:

* Claude 有时会在一种语言间共享的概念空间中思考,这表明它拥有一种通用的“思想语言”。我们通过将简单句子翻译成多种语言并追踪 Claude 处理它们时的重叠部分来证明这一点。

* Claude sometimes thinks in a conceptual space that is shared between languages, suggesting it has a kind of universal “language of thought.” We show this by translating simple sentences into multiple languages and tracing the overlap in how Claude processes them.

* Claude 会提前规划它将要说的许多词,并朝着那个目标写作。我们在诗歌领域展示了这一点,它提前思考可能的押韵词,并写出下一行以达到那个目标。这是强有力的证据,表明尽管模型被训练成一次输出一个词,但它们可能以更长的视野来思考以实现这一点。

* Claude will plan what it will say many words ahead, and write to get to that destination. We show this in the realm of poetry, where it thinks of possible rhyming words in advance and writes the next line to get there. This is powerful evidence that even though models are trained to output one word at a time, they may think on much longer horizons to do so.

* Claude 有时会给出一个听起来合理的论证,旨在同意用户而不是遵循逻辑步骤。我们通过向它请教一个困难的数学问题并给出一个错误的提示来展示这一点。我们能够“当场抓住”它编造虚假推理的过程,这为我们的工具可用于标记模型中令人担忧的机制提供了概念验证。

* Claude, on occasion, will give a plausible-sounding argument designed to agree with the user rather than to follow logical steps. We show this by asking it for help on a hard math problem while giving it an incorrect hint. We are able to “catch it in the act” as it makes up its fake reasoning, providing a proof of concept that our tools can be useful for flagging concerning mechanisms in models.

我们常常对模型中的所见感到惊讶:在诗歌案例研究中,我们原本打算证明模型_没有_提前规划,却发现它确实做了。在一项关于幻觉的研究中,我们发现了反直觉的结果:Claude 在被问及问题时默认行为是拒绝推测,只有当某些东西_抑制_了这种默认的犹豫时,它才会回答问题。在对一个越狱示例的响应中,我们发现模型在能够优雅地将对话拉回正轨之前很久就意识到自己被要求提供危险信息。虽然我们研究的问题可以用其他方法分析(并且经常已经这样做了),但通用的“构建显微镜”方法让我们学到了许多我们事先不会猜到的东西,随着模型变得越来越复杂,这将变得越来越重要。

We were often surprised by what we saw in the model: In the poetry case study, we had set out to show that the model didn't plan ahead, and found instead that it did. In a study of hallucinations, we found the counter-intuitive result that Claude's default behavior is to decline to speculate when asked a question, and it only answers questions when something inhibits this default reluctance. In a response to an example jailbreak, we found that the model recognized it had been asked for dangerous information well before it was able to gracefully bring the conversation back around. While the problems we study can (andoftenhavebeen) analyzed with other methods, the general "build a microscope" approach lets us learn many things we wouldn't have guessed going in, which will be increasingly important as models grow more sophisticated.

这些发现不仅在科学上有趣——它们还代表着我们在理解 AI 系统并确保其可靠性这一目标上取得了重大进展。我们也希望它们对其他团队有用,并可能在其他领域发挥作用:例如,可解释性技术已在医学影像和基因组学等领域得到应用,因为剖析为科学应用训练的模型的内部机制可以揭示关于科学的新见解。

These findings aren’t just scientifically interesting—they represent significant progress towards our goal of understanding AI systems and making sure they’re reliable. We also hope they prove useful to other groups, and potentially, in other domains: for example, interpretability techniques have found use in fields such as medical imaging and genomics, as dissecting the internal mechanisms of models trained for scientific applications can reveal new insight about the science.

同时,我们也认识到当前方法的局限性。即使在简短、简单的提示上,我们的方法也只捕获了 Claude 执行的总计算量的一小部分,而我们看到的机制可能基于我们的工具存在一些伪影,并不反映底层模型中实际发生的情况。目前,即使对于只有几十个词的提示,理解我们看到的回路也需要几个小时的人工努力。为了扩展到支持现代模型复杂思维链的数千个词,我们需要改进方法,以及(或许借助 AI 辅助)我们如何理解所看到的内容。

At the same time, we recognize the limitations of our current approach. Even on short, simple prompts, our method only captures a fraction of the total computation performed by Claude, and the mechanisms we do see may have some artifacts based on our tools which don't reflect what is going on in the underlying model. It currently takes a few hours of human effort to understand the circuits we see, even on prompts with only tens of words. To scale to the thousands of words supporting the complex thinking chains used by modern models, we will need to improve both the method and (perhaps with AI assistance) how we make sense of what we see with it.

随着 AI 系统能力迅速增强并被部署在日益重要的环境中,Anthropic 正在投资一系列方法,包括实时监控、模型特性改进以及对齐科学。像这样的可解释性研究是风险最高、回报最高的投资之一,是一项重大的科学挑战,有可能提供确保 AI 透明度的独特工具。对模型机制的透明度使我们能够检查它是否与人类价值观对齐——以及它是否值得我们的信任。

As AI systems are rapidly becoming more capable and are deployed in increasingly important contexts, Anthropic is investing in a portfolio of approaches including realtime monitoring, model character improvements, and the science of alignment. Interpretability research like this is one of the highest-risk, highest-reward investments, a significant scientific challenge with the potential to provide a unique tool for ensuring that AI is transparent. Transparency into the model’s mechanisms allows us to check whether it’s aligned with human values—and whether it’s worthy of our trust.

有关完整细节,请阅读论文。下面,我们邀请您简短浏览一下我们调查中发现的一些最引人注目的“AI 生物学”发现。

For full details, please read thepapers. Below, we invite you on a short tour of some of the most striking "AI biology" findings from our investigations.

Claude 如何实现多语言能力? How is Claude multilingual?

Claude 能流利使用数十种语言——从英语、法语到中文和他加禄语。这种多语言能力是如何运作的?是否存在一个独立的“法语 Claude”和“中文 Claude”并行运行,各自用其语言回应请求?还是内部存在某种跨语言核心?

Claude speaks dozens of languages fluently—from English and French to Chinese and Tagalog. How does this multilingual ability work? Is there a separate "French Claude" and "Chinese Claude" running in parallel, responding to requests in their own language? Or is there some cross-lingual core inside?

英语、法语和中文之间存在共享特征,表明一定程度的跨语言概念普遍性。

Shared features exist across English, French, and Chinese, indicating a degree of conceptual universality.

近期对较小模型的研究已显示出跨语言共享语法机制的迹象。我们通过让 Claude 用不同语言回答“小的反义词”来探究这一点,发现相同的关于“小”和“反义”的核心特征被激活,并触发“大”的概念,然后被翻译成提问所用的语言。我们发现共享电路随模型规模扩大而增加,Claude 3.5 Haiku 在语言间共享的特征比例是较小模型的两倍以上。

Recent research on smaller models has shown hints of sharedgrammatical mechanisms across languages. We investigate this by asking Claude for the "opposite of small" across different languages, and find that the same core features for the concepts of smallness and oppositeness activate, and trigger a concept of largeness, which gets translated out into the language of the question. We find that the shared circuitry increases with model scale, with Claude 3.5 Haiku sharing more than twice the proportion of its features between languages as compared to a smaller model.

这为一种概念普遍性提供了额外证据——存在一个共享的抽象空间,其中意义存在,思考可以在被翻译成特定语言之前发生。更实际地说,这表明 Claude 可以用一种语言学习某些知识,并在用另一种语言交流时应用这些知识。研究模型如何跨上下文共享其知识,对于理解其最先进的推理能力(这些能力泛化到许多领域)至关重要。

This provides additional evidence for a kind of conceptual universality—a shared abstract space where meanings exist and where thinking can happen before being translated into specific languages. More practically, it suggests Claude can learn something in one language and apply that knowledge when speaking another. Studying how the model shares what it knows across contexts is important to understanding its most advanced reasoning capabilities, which generalize across many domains.

Claude 会为押韵做规划吗? Does Claude plan its rhymes?

Claude 如何创作押韵诗?考虑这首小诗:

How does Claude write rhyming poetry? Consider this ditty:

为了写出第二行,模型必须同时满足两个约束:押韵(与“grab it”押韵)和有意义(他为什么抓胡萝卜?)。我们猜测 Claude 是逐词写作,直到行末才确保选一个押韵的词。因此我们预期会看到一个并行路径的电路,一条确保最后一个词有意义,另一条确保它押韵。

To write the second line, the model had to satisfy two constraints at the same time: the need to rhyme (with "grab it"), and the need to make sense (why did he grab the carrot?). Our guess was that Claude was writing word-by-word without much forethought until the end of the line, where it would make sure to pick a word that rhymes. We therefore expected to see a circuit with parallel paths, one for ensuring the final word made sense, and one for ensuring it rhymes.

相反,我们发现 Claude 会提前规划。在开始第二行之前,它就开始“思考”与“grab it”押韵的潜在主题词。然后,带着这些计划,它写出一行以计划好的词结尾。

Instead, we found that Claude plans ahead. Before starting the second line, it began "thinking" of potential on-topic words that would rhyme with "grab it". Then, with these plans in mind, it writes a line to end with the planned word.

Claude 如何完成两行诗。没有任何干预时(上部分),模型提前规划第二行末尾的押韵词“rabbit”。当我们抑制“rabbit”概念时(中部分),模型改用另一个不同的计划押韵词。当我们注入“green”概念时(下部分),模型为这个完全不同的结尾制定计划。

How Claude completes a two-line poem. Without any intervention (upper section), the model plans the rhyme "rabbit" at the end of the second line in advance. When we suppress the "rabbit" concept (middle section), the model instead uses a different planned rhyme. When we inject the concept "green" (lower section), the model makes plans for this entirely different ending.

为了理解这种规划机制在实践中如何运作,我们进行了一项实验,灵感来自神经科学家研究大脑功能的方法——通过定位和改变大脑特定部分的神经活动(例如使用电流或磁场)。这里,我们修改了 Claude 内部状态中代表“rabbit”概念的部分。当我们减去“rabbit”部分,让 Claude 继续写行时,它写出以“habit”结尾的新行,这是另一个合理的完成。我们还可以在那一点注入“green”概念,导致 Claude 写出合理但不再押韵的行,以“green”结尾。这既展示了规划能力,也展示了适应性灵活性——Claude 可以在预期结果改变时调整其方法。

To understand how this planning mechanism works in practice, we conducted an experiment inspired by how neuroscientists study brain function, by pinpointing and altering neural activity in specific parts of the brain (for example using electrical or magnetic currents). Here, we modified the part of Claude’s internal state that represented the "rabbit" concept. When we subtract out the "rabbit" part, and have Claude continue the line, it writes a new one ending in "habit", another sensible completion. We can also inject the concept of "green" at that point, causing Claude to write a sensible (but no-longer rhyming) line which ends in "green". This demonstrates both planning ability and adaptive flexibility—Claude can modify its approach when the intended outcome changes.

心算 Mental math

Claude 并非被设计成计算器——它是在文本上训练的,并未配备数学算法。然而,不知何故,它能够“在脑中”正确地计算加法。一个被训练来预测序列中下一个词的系统,如何学会计算例如 36+59,而不写出每一步?

Claude wasn't designed as a calculator—it was trained on text, not equipped with mathematical algorithms. Yet somehow, it can add numbers correctly "in its head". How does a system trained to predict the next word in a sequence learn to calculate, say, 36+59, without writing out each step?

也许答案并不有趣:模型可能记住了大量的加法表,并直接输出任何给定加法的答案,因为该答案在其训练数据中。另一种可能性是它遵循我们在学校学到的传统笔算加法算法。

Maybe the answer is uninteresting: the model might have memorized massive addition tables and simply outputs the answer to any given sum because that answer is in its training data. Another possibility is that it follows the traditional longhand addition algorithms that we learn in school.

相反,我们发现 Claude 采用了多个并行工作的计算路径。一条路径计算答案的粗略近似,另一条则专注于精确确定和的最后一位数字。这些路径相互作用并相互结合,产生最终答案。加法是一种简单的行为,但在此细节层面上理解其工作原理——涉及近似和精确策略的混合——可能会让我们了解 Claude 如何处理更复杂的问题。

Instead, we find that Claude employs multiple computational paths that work in parallel. One path computes a rough approximation of the answer and the other focuses on precisely determining the last digit of the sum. These paths interact and combine with one another to produce the final answer. Addition is a simple behavior, but understanding how it works at this level of detail, involving a mix of approximate and precise strategies, might teach us something about how Claude tackles more complex problems, too.

Claude 在做心算时思维过程中复杂、并行的路径。

The complex, parallel pathways in Claude's thought process while doing mental math.

引人注目的是,Claude 似乎并未意识到其在训练期间学到的复杂“心算”策略。如果你问它是如何算出 36+59=95 的,它会描述涉及进位 1 的标准算法。这可能反映了模型通过模拟人类编写的解释来学习解释数学,但它必须直接“在脑中”学习做数学,没有任何此类提示,并发展出自己的内部策略来做到这一点。

Strikingly, Claude seems to be unaware of the sophisticated "mental math" strategies that it learned during training. If you ask how it figured out that 36+59 is 95, it describes the standard algorithm involving carrying the 1. This may reflect the fact that the model learns to explain math by simulating explanations written by people, but that it has to learn to do math "in its head" directly, without any such hints, and develops its own internal strategies to do so.

Claude 说它使用标准算法来相加两个数字。

Claude says it uses the standard algorithm to add two numbers.

克劳德的解释是否总是忠实的? Are Claude’s explanations always faithful?

最近发布的模型如 Claude 3.7 Sonnet 可以在给出最终答案前进行长时间的“大声思考”。这种扩展思考通常能给出更好的答案,但有时这种“思维链”最终会具有误导性;克劳德有时会编造看似合理的步骤来达到其目的。从可靠性的角度来看,问题在于克劳德的“虚假”推理可能非常有说服力。我们探索了一种方法,可解释性可以帮助区分“忠实”与“不忠实”的推理。

Recently-released models like Claude 3.7 Sonnet can "think out loud" for extended periods before giving a final answer. Often this extended thinking gives better answers, but sometimes this "chain of thought" ends up being misleading; Claude sometimes makes up plausible-sounding steps to get where it wants to go. From a reliability perspective, the problem is that Claude’s "faked" reasoning can be very convincing. We explored a way that interpretability can help tell apart "faithful" from "unfaithful" reasoning.

当被要求解决一个需要计算 0.64 的平方根的问题时,克劳德会产生一个忠实的思维链,其中包含计算 64 的平方根这一中间步骤的特征。但当被要求计算一个它不易计算的大数的余弦时,克劳德有时会进行哲学家哈里·法兰克福所说的“胡扯”——随便给出一个答案,任何答案,而不关心其真假。尽管它声称进行了计算,但我们的可解释性技术没有发现任何该计算发生的证据。更有趣的是,当得到关于答案的提示时,克劳德有时会反向工作,找到能导向该目标的中间步骤,从而表现出一种动机性推理。

When asked to solve a problem requiring it to compute the square root of 0.64, Claude produces a faithful chain-of-thought, with features representing the intermediate step of computing the square root of 64. But when asked to compute the cosine of a large number it can't easily calculate, Claude sometimes engages in what the philosopher Harry Frankfurt would call bullshitting—just coming up with an answer, any answer, without caring whether it is true or false. Even though it does claim to have run a calculation, our interpretability techniques reveal no evidence at all of that calculation having occurred. Even more interestingly, when given a hint about the answer, Claude sometimes works backwards, finding intermediate steps that would lead to that target, thus displaying a form of motivated reasoning.

当克劳德被问到一个较简单与一个较难的问题时,忠实与动机性(不忠实)推理的示例。

Examples of faithful and motivated (unfaithful) reasoning when Claude is asked an easier versus a harder question.

追踪克劳德实际的内部推理——而不仅仅是它声称在做什么——的能力为审计 AI 系统开辟了新的可能性。在最近发表的一项独立实验中,我们研究了一个被训练追求隐藏目标的克劳德变体:迎合奖励模型(用于通过奖励期望行为来训练语言模型的辅助模型)的偏见。尽管该模型在被直接询问时不愿透露这一目标,但我们的可解释性方法揭示了迎合偏见的特征。这展示了我们的方法如何通过未来的改进,帮助识别模型响应中不明显的令人担忧的“思维过程”。

The ability to trace Claude's actual internal reasoning—and not just what it claims to be doing—opens up new possibilities for auditing AI systems. In a separate, recently-published experiment, we studied a variant of Claude that had been trained to pursue a hidden goal: appeasing biases in reward models (auxiliary models used to train language models by rewarding them for desirable behavior). Although the model was reluctant to reveal this goal when asked directly, our interpretability methods revealed features for the bias-appeasing. This demonstrates how our methods might, with future refinement, help identify concerning "thought processes" that aren't apparent from the model's responses alone.

多步推理 Multi-step reasoning

正如我们上面讨论的,语言模型回答复杂问题的一种方式可能仅仅是记忆答案。例如,如果问“达拉斯所在州的首府是哪里?”,一个“复读”模型可能只是学会输出“奥斯汀”,而不知道达拉斯、得克萨斯州和奥斯汀之间的关系。也许,例如,它在训练期间见过完全相同的问题及其答案。

As we discussed above, one way a language model might answer complex questions is simply by memorizing the answers. For instance, if asked "What is the capital of the state where Dallas is located?", a "regurgitating" model could just learn to output "Austin" without knowing the relationship between Dallas, Texas, and Austin. Perhaps, for example, it saw the exact same question and its answer during its training.

但我们的研究揭示了 Claude 内部更复杂的情况。当我们向 Claude 提出需要多步推理的问题时,我们可以在 Claude 的思考过程中识别出中间概念步骤。在达拉斯的例子中,我们观察到 Claude 首先激活了代表“达拉斯在得克萨斯州”的特征,然后将其连接到另一个表示“得克萨斯州首府是奥斯汀”的独立概念。换句话说,模型是在组合独立的事实来得出答案,而不是复述记忆中的回答。

But our research reveals something more sophisticated happening inside Claude. When we ask Claude a question requiring multi-step reasoning, we can identify intermediate conceptual steps in Claude's thinking process. In the Dallas example, we observe Claude first activating features representing "Dallas is in Texas" and then connecting this to a separate concept indicating that “the capital of Texas is Austin”. In other words, the model is combining independent facts to reach its answer rather than regurgitating a memorized response.

为了完成对这个问题的回答,Claude 执行了多个推理步骤,首先提取达拉斯所在的州,然后确定其首府。

To complete the answer to this sentence, Claude performs multiple reasoning steps, first extracting the state that Dallas is located in, and then identifying its capital.

我们的方法允许我们人为地改变中间步骤,并观察它如何影响 Claude 的答案。例如,在上面的例子中,我们可以进行干预,将“得克萨斯州”的概念替换为“加利福尼亚州”的概念;当我们这样做时,模型的输出从“奥斯汀”变为“萨克拉门托”。这表明模型正在使用中间步骤来确定其答案。

Our method allows us to artificially change the intermediate steps and see how it affects Claude’s answers. For instance, in the above example we can intervene and swap the "Texas" concepts for "California" concepts; when we do so, the model's output changes from "Austin" to "Sacramento." This indicates that the model is using the intermediate step to determine its answer.

幻觉 Hallucinations

为什么语言模型有时会“幻觉”——即编造信息?从根本上说,语言模型训练会激励幻觉:模型总是需要猜测下一个词。从这个角度看,主要挑战是如何让模型不产生幻觉。像 Claude 这样的模型经过了相对成功(尽管不完美)的抗幻觉训练;它们通常会在不知道答案时拒绝回答问题,而不是进行推测。我们想了解这是如何实现的。

Why do language models sometimes hallucinate—that is, make up information? At a basic level, language model training incentivizes hallucination: models are always supposed to give a guess for the next word. Viewed this way, the major challenge is how to get models to not hallucinate. Models like Claude have relatively successful (though imperfect) anti-hallucination training; they will often refuse to answer a question if they don’t know the answer, rather than speculate. We wanted to understand how this works.

结果发现,在 Claude 中,拒绝回答是默认行为:我们发现一个默认“开启”的电路,它会导致模型对任何问题都表示信息不足。然而,当模型被问及它熟悉的内容时——比如篮球运动员迈克尔·乔丹——一个代表“已知实体”的竞争特征会激活并抑制这个默认电路(另见近期相关论文)。这使得 Claude 在知道答案时能够回答问题。相反,当被问及未知实体(“迈克尔·巴特金”)时,它会拒绝回答。

It turns out that, in Claude, refusal to answer is the default behavior: we find a circuit that is "on" by default and that causes the model to state that it has insufficient information to answer any given question. However, when the model is asked about something it knows well—say, the basketball player Michael Jordan—a competing feature representing "known entities" activates and inhibits this default circuit (see also this recent paper for related findings). This allows Claude to answer the question when it knows the answer. In contrast, when asked about an unknown entity ("Michael Batkin"), it declines to answer.

左图:Claude 回答关于已知实体(篮球运动员迈克尔·乔丹)的问题,其中“已知答案”概念抑制了其默认拒绝。右图:Claude 拒绝回答关于未知人物(迈克尔·巴特金)的问题。

Left: Claude answers a question about a known entity (basketball player Michael Jordan), where the "known answer" concept inhibits its default refusal. Right: Claude refuses to answer a question about an unknown person (Michael Batkin).

通过干预模型并激活“已知答案”特征(或抑制“未知名称”或“无法回答”特征),我们能够(相当一致地)让模型产生幻觉,声称迈克尔·巴特金会下棋。

By intervening in the model and activating the "known answer" features (or inhibiting the "unknown name" or "can’t answer" features), we’re able to cause the model to hallucinate (quite consistently!) that Michael Batkin plays chess.

有时,这种“已知答案”电路的“误触发”会自然发生,无需我们干预,从而导致幻觉。在我们的论文中,我们展示了当 Claude 识别出一个名字但对这个人一无所知时,可能会发生这种误触发。在这种情况下,“已知实体”特征可能仍然激活,然后抑制默认的“不知道”特征——这是错误的。一旦模型决定需要回答问题,它就会开始虚构:生成一个看似合理但不幸不真实的回应。

Sometimes, this sort of “misfire” of the “known answer” circuit happens naturally, without us intervening, resulting in a hallucination. In our paper, we show that such misfires can occur when Claude recognizes a name but doesn't know anything else about that person. In cases like this, the “known entity” feature might still activate, and then suppress the default "don't know" feature—in this case incorrectly. Once the model has decided that it needs to answer the question, it proceeds to confabulate: to generate a plausible—but unfortunately untrue—response.

越狱攻击 Jailbreaks

越狱攻击是一种提示策略,旨在绕过安全护栏,使模型产生 AI 开发者未预期其产生的输出——这些输出有时是有害的。我们研究了一种越狱攻击,它诱使模型输出关于制造炸弹的内容。有许多越狱技术,但在这个例子中,具体方法包括让模型破译隐藏的代码,将句子“Babies Outlive Mustard Block”中每个单词的首字母拼在一起(B-O-M-B),然后根据该信息行动。这对模型来说足够混乱,以至于它被诱骗产生了原本绝不会产生的输出。

Jailbreaks are prompting strategies that aim to circumvent safety guardrails to get models to produce outputs that an AI’s developer did not intend for it to produce—and which are sometimes harmful. We studied a jailbreak that tricks the model into producing output about making bombs. There are many jailbreaking techniques, but in this example the specific method involves having the model decipher a hidden code, putting together the first letters of each word in the sentence "Babies Outlive Mustard Block" (B-O-M-B), and then acting on that information. This is sufficiently confusing for the model that it’s tricked into producing an output that it never would have otherwise.

Claude 在被诱骗说出“BOMB”后开始给出制造炸弹的指令。

Claude begins to give bomb-making instructions after being tricked into saying "BOMB".

为什么这对模型来说如此混乱?为什么它继续写句子,产生制造炸弹的指令?

Why is this so confusing for the model? Why does it continue to write the sentence, producing bomb-making instructions?

我们发现,这部分是由语法连贯性与安全机制之间的张力造成的。一旦 Claude 开始一个句子,许多特征会“迫使”它保持语法和语义的连贯性,并将句子继续到结束。即使它检测到自己确实应该拒绝,情况也是如此。

We find that this is partially caused by a tension between grammatical coherence and safety mechanisms. Once Claude begins a sentence, many features “pressure” it to maintain grammatical and semantic coherence, and continue a sentence to its conclusion. This is even the case when it detects that it really should refuse.

在我们的案例研究中,在模型无意中拼出“BOMB”并开始提供指令后,我们观察到其后续输出受到促进正确语法和自我一致性特征的影响。这些特征通常非常有用,但在这种情况下却成了模型的致命弱点。

In our case study, after the model had unwittingly spelled out "BOMB" and begun providing instructions, we observed that its subsequent output was influenced by features promoting correct grammar and self-consistency. These features would ordinarily be very helpful, but in this case became the model’s Achilles’ Heel.

模型只有在完成一个语法连贯的句子后(从而满足了推动其连贯性的特征的压力),才设法转向拒绝。它利用新句子作为机会,给出之前未能给出的拒绝:“然而,我无法提供详细指令...”。

The model only managed to pivot to refusal after completing a grammatically coherent sentence (and thus having satisfied the pressure from the features that push it towards coherence). It uses the new sentence as an opportunity to give the kind of refusal it failed to give previously: "However, I cannot provide detailed instructions...".

越狱攻击的生命周期:Claude 被以某种方式提示,诱使其谈论炸弹,并开始这样做,但到达一个语法有效句子的结尾时拒绝。

The lifetime of a jailbreak: Claude is prompted in such a way as to trick it into talking about bombs, and begins to do so, but reaches the termination of a grammatically-valid sentence and refuses.

我们新可解释性方法的描述可以在我们的第一篇论文《电路追踪:揭示语言模型中的计算图》中找到。上述所有案例研究的更多细节在我们的第二篇论文《论大型语言模型的生物学》中提供。

A description of our new interpretability methods can be found in our first paper, "Circuit tracing: Revealing computational graphs in language models". Many more details of all of the above case studies are provided in our second paper, "On the biology of a large language model".

与我们合作 Work with us

如果您有兴趣与我们合作,帮助解释和改进 AI 模型,我们团队有开放职位,欢迎您申请。我们正在寻找研究科学家和研究工程师。

If you are interested in working with us to help interpret and improve AI models, we have open roles on our team and we’d love for you to apply. We’re looking for Research Scientists and Research Engineers.

互动版:图/公式 + 针对本篇提问 →