语言模型可以解释语言模型中的神经元

Language models can explain neurons in language models

OpenAI OpenAI · OpenAI · 2023-05-09 · OpenAI Research ↗

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

语言模型可以解释语言模型中的神经元 | OpenAI 阅读论文(在新窗口中打开)查看神经元(在新窗口中打开)查看代码和数据集(在新窗口中打开)我们使用 GPT-4 自动为大型语言模型中神经元的行为编写解释,并对这些解释进行评分。我们发布了这些(不完美的)解释和 GPT-2 中每个神经元的评分数据集。

Language models can explain neurons in language models | OpenAI Read paper(opens in a new window)View neurons(opens in a new window)View code and dataset(opens in a new window) We use GPT‑4 to automatically write explanations for the behavior of neurons in large language models and to score those explanations. We release a dataset of these (imperfect) explanations and scores for every neuron in GPT‑2.

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 8)

全文 · Full text(逐段中英对照)

概述 Overview

语言模型可以解释语言模型中的神经元 | OpenAI

Language models can explain neurons in language models | OpenAI

阅读论文(在新窗口打开)查看神经元(在新窗口打开)查看代码和数据集(在新窗口打开)

Read paper(opens in a new window)View neurons(opens in a new window)View code and dataset(opens in a new window)

我们使用 GPT-4 自动编写大型语言模型中神经元行为的解释,并对这些解释进行评分。我们发布了 GPT-2 中每个神经元的这些(不完美的)解释和评分数据集。

We use GPT‑4 to automatically write explanations for the behavior of neurons in large language models and to score those explanations. We release a dataset of these (imperfect) explanations and scores for every neuron in GPT‑2.

语言模型已经变得更有能力且部署更广泛,但我们对其内部工作机制的理解仍然非常有限。例如,可能很难从其输出中检测它们是否使用了有偏见的启发式方法或进行欺骗。可解释性研究旨在通过查看模型内部来揭示额外信息。

Language models have become more capable and more broadly deployed, but our understanding of how they work internally is still very limited. For example, it might be difficult to detect from their outputs whether they use biased heuristics or engage in deception. Interpretability research aims to uncover additional information by looking inside the model.

可解释性研究的一种简单方法是首先理解各个组件(神经元和注意力头)在做什么。传统上,这需要人类手动检查神经元以找出它们代表的数据特征。这个过程难以扩展:很难将其应用于具有数百亿参数规模的神经网络。我们提出了一种自动化过程,使用 GPT-4 生成和评分神经元行为的自然语言解释,并将其应用于另一个语言模型中的神经元。

One simple approach to interpretability research is to first understand what the individual components (neurons and attention heads) are doing. This has traditionally required humans to manually⁠(opens in a new window)inspect⁠neurons⁠(opens in a new window) to figure out what features of the data they represent. This process doesn’t scale well: it’s hard to apply it to neural networks with tens or hundreds of billions of parameters. We propose an automated process that uses GPT‑4 to produce and score natural language explanations of neuron behavior and apply it to neurons in another language model.

这项工作是我们对齐研究方法第三支柱的一部分:我们希望自动化对齐研究工作本身。这种方法的一个有前景的方面是它随着 AI 发展的步伐而扩展。随着未来模型作为助手变得越来越智能和有用,我们将找到更好的解释。

This work is part of the third pillar of our approach to alignment research⁠: we want to automate the alignment research work itself. A promising aspect of this approach is that it scales with the pace of AI development. As future models become increasingly intelligent and helpful as assistants, we will find better explanations.

工作原理 How it works

我们的方法包括对每个神经元执行 3 个步骤。

Our methodology consists of running 3 steps on every neuron.

步骤 1:使用 GPT-4 生成解释 Step 1: Generate explanation using GPT-4

给定一个 GPT-2 神经元,通过向 GPT-4 展示相关文本序列和激活值,生成对其行为的解释。

Given a GPT-2 neuron, generate an explanation of its behavior by showing relevant text sequences and activations to GPT-4.

步骤 2:使用 GPT-4 进行模拟 Step 2: Simulate using GPT-4

再次使用 GPT-4 模拟一个因该解释而激活的神经元会做什么。

Simulate what a neuron that fired for the explanation would do, again using GPT-4

第三步:比较 Step 3: Compare

根据模拟激活与真实激活的匹配程度对解释进行评分。

Score the explanation based on how well the simulated activations match the real activations

我们的发现 What we found

使用我们的评分方法,我们可以开始衡量我们的技术对网络不同部分的效果,并尝试改进对当前解释不佳的部分的技术。例如,我们的技术对较大的模型效果不佳,可能是因为后面的层更难解释。

Using our scoring methodology, we can start to measure how well our techniques work for different parts of the network and try to improve the technique for parts that are currently poorly explained. For example, our technique works poorly for larger models, possibly because later layers are harder to explain.

尽管绝大多数解释得分较低,但我们相信现在可以利用机器学习技术进一步提高我们生成解释的能力。例如,我们发现可以通过以下方式提高得分:

Although the vast majority of our explanations score poorly, we believe we can now use ML techniques to further improve our ability to produce explanations. For example, we found we were able to improve scores by:

* 迭代解释。 我们可以通过让 GPT‑4 提出可能的反例,然后根据激活值修改解释来提高得分。

* Iterating on explanations. We can increase scores by asking GPT‑4 to come up with possible counterexamples, then revising explanations in light of their activations.

* 使用更大的模型来提供解释。 平均得分随着解释器模型能力的增强而提高。然而,即使是 GPT‑4 给出的解释也比人类差,这表明还有改进空间。

* Using larger models to give explanations. The average score goes up as the explainer model’s capabilities increase. However, even GPT‑4 gives worse explanations than humans, suggesting room for improvement.

* 改变被解释模型的架构。 使用不同激活函数训练模型可以提高解释得分。

* Changing the architecture of the explained model. Training models with different activation functions improved explanation scores.

我们正在开源我们的数据集和可视化工具,用于 GPT‑4 对 GPT‑2 中所有 307,200 个神经元编写的解释,以及使用 OpenAI API 上公开可用的模型进行解释和评分的代码。我们希望研究社区能够开发新技术来生成更高得分的解释,以及更好的工具来利用解释探索 GPT‑2。

We are open-sourcing our datasets and visualization tools for GPT‑4‑written explanations of all 307,200 neurons in GPT‑2, as well as code for explanation and scoring using publicly available models⁠(opens in a new window) on the OpenAI API. We hope the research community will develop new techniques for generating higher-scoring explanations and better tools for exploring GPT‑2 using explanations.

我们发现了超过 1,000 个神经元,其解释得分至少为 0.8,这意味着根据 GPT‑4,这些解释解释了神经元大部分最高激活行为。这些解释良好的神经元大多并不十分有趣。然而,我们也发现许多有趣的神经元是 GPT‑4 无法理解的。我们希望随着解释的改进,我们能够迅速揭示对模型计算的有趣定性理解。

We found over 1,000 neurons with explanations that scored at least 0.8, meaning that according to GPT‑4 they account for most of the neuron’s top-activating behavior. Most of these well-explained neurons are not very interesting. However, we also found many interesting neurons that GPT‑4 didn't understand. We hope as explanations improve we may be able to rapidly uncover interesting qualitative understanding of model computations.

4 个样本中的第 1 个 Sample 1 of 4

我们的许多读者可能知道,日本消费者非常喜欢独特且富有创意的奇巧产品和口味。但现在,雀巢日本推出了一款不仅可称为新口味,更可称为奇巧新“物种”的产品。

Many of our readers may be aware that Japanese consumers are quite fond of unique and creative Kit Kat products and flavors. But now, Nestle Japan has come out with what could be described as not just a new flavor but a new "species" of Kit Kat.

“大写字母‘K’后接各种字母组合”

“uppercase ‘K’ followed by various combinations of letters”

“与品牌名和商业相关的单词和短语的部分”

“parts of words and phrases related to brand names and businesses”

神经元在各层间激活,更高层更抽象。

Neurons activating across layers, higher layers are more abstract.

展望 Outlook

我们的方法目前有许多局限性,我们希望未来的工作能够解决这些问题。

Our method currently has many limitations⁠(opens in a new window), which we hope can be addressed in future work.

* 我们专注于简短的自然语言解释,但神经元可能具有非常复杂的行为,无法简洁地描述。例如,神经元可能高度多义(表示许多不同的概念),或者可能表示人类不理解或没有词语描述的单一概念。

* We focused on short natural language explanations, but neurons may have very complex behavior that is impossible to describe succinctly. For example, neurons could be highly polysemantic (representing many distinct concepts) or could represent single concepts that humans don't understand or have words for.

* 我们最终希望自动发现并解释实现复杂行为的整个神经回路,其中神经元和注意力头协同工作。我们当前的方法仅将神经元行为解释为原始文本输入的函数,而没有说明其下游影响。例如,一个在句号上激活的神经元可能表示下一个单词应以大写字母开头,或者正在递增句子计数器。

* We want to eventually automatically find and explain entire neural circuits⁠(opens in a new window) implementing complex behaviors, with neurons and attention heads working together. Our current method only explains neuron behavior as a function of the original text input, without saying anything about its downstream effects. For example, a neuron that activates on periods could be indicating the next word should start with a capital letter, or be incrementing a sentence counter.

* 我们解释了神经元的行为,但没有尝试解释产生该行为的机制。这意味着即使得分很高的解释在分布外的文本上也可能表现很差,因为它们只是描述了一种相关性。

* We explained the behavior of neurons without attempting to explain the mechanisms that produce that behavior. This means that even high-scoring explanations could do very poorly on out-of-distribution texts, since they are simply describing a correlation.

* 我们的整体过程计算量相当大。

* Our overall procedure is quite compute intensive.

我们对方法的扩展和推广感到兴奋。最终,我们希望像可解释性研究者一样,使用模型来形成、测试和迭代完全通用的假设。

We are excited about extensions and generalizations of our approach. Ultimately, we would like to use models to form, test, and iterate on fully general hypotheses just as an interpretability researcher would.

最终,我们希望解释我们最大的模型,作为在部署前后检测对齐和安全问题的一种方式。然而,在这些技术能够揭示像不诚实这样的行为之前,我们还有很长的路要走。

Eventually we want to interpret our largest models as a way to detect alignment and safety problems before and after deployment. However, we still have a long way to go before these techniques can surface behaviors like dishonesty.

互动版:图/公式 + 针对本篇提问 →