Should Developers Care about Interpretability?
打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→可解释性和导引是如何工作的?承诺有哪些应用?1. 捕捉无法用言语描述的风格 2. 减少 RLHF 需求 3. 记住用户偏好与用户请求 4. 廉价、快速、可复现的分类 缺点有哪些?1. 使模型‘偏离分布’ 2. 不理解特征的作用 3. 激活其他特征和电路 结论 Thariq Shihipar - 2024 年 11 月 4 日 · 6 分钟阅读 今年 LLM 研究中最大的突破之一就是可解释性——即理解 LLM 在“思考”什么的能力。
How does Interpretability and Steering work?The PromiseWhat are the applications?1. Capturing style that cannot be described in words2. Less need for RLHF3. Remembering User Preferences vs User Requests4. Cheap, Fast, Reproducible ClassificationWhat are the downsides?1. Moving the model ‘out of distribution’2. Not understanding what a feature does3. Activating other features & circuitsConclusion Thariq Shihipar - 4 November 2024 · 6 min read Arguably the biggest breakthrough in LLM research this year has been in interpretability- the ability to understand what a LLM is “thinking”.
可解释性与引导如何工作?前景如何?应用有哪些?1. 捕捉无法用语言描述的风格 2. 减少对 RLHF 的需求 3. 记住用户偏好与用户请求 4. 廉价、快速、可复现的分类 有哪些缺点?1. 将模型移出分布 2. 不理解特征的作用 3. 激活其他特征与回路 结论
How does Interpretability and Steering work?The PromiseWhat are the applications?1. Capturing style that cannot be described in words2. Less need for RLHF3. Remembering User Preferences vs User Requests4. Cheap, Fast, Reproducible ClassificationWhat are the downsides?1. Moving the model ‘out of distribution’2. Not understanding what a feature does3. Activating other features & circuitsConclusion
Thariq Shihipar - 2024 年 11 月 4 日 · 6 分钟阅读
Thariq Shihipar - 4 November 2024 · 6 min read
可以说,今年 LLM 研究中最大的突破在于可解释性——即理解 LLM“思考”内容的能力。
Arguably the biggest breakthrough in LLM research this year has been in interpretability- the ability to understand what a LLM is “thinking”.
最著名的例子是 Anthropic 的 Golden Gate Claude,尽管这项工作并不局限于文本,研究人员还在研究图像、语音甚至蛋白质模型。
The most famous example is Anthropic’s Golden Gate Claude, though this work isn’t limited to text, researchers are also working on images, voice and even protein models.
但是,虽然可解释性最常在研究与 AI 安全的背景下被讨论,它也为开发者提供了对其模型更细粒度的控制与可靠性的承诺。
But while interpretability is most often discussed in the context of research & AI safety, it also offers a promise to developers of more fine-grained control and reliability from their models.
最好的基础解释来自《金门大桥 Claude》论文,它既详尽又易读,但我已尽力从 AI 开发者的角度总结可解释性和引导的工作。
The best foundational explanation is in the Golden Gate Claude paper which is both thorough and easy to read, but I have tried my best to summarize interpretability and steering work from the view of an AI dev.
* 你可以将 LLM 分解为一组‘特征’,这些特征描述不同的概念,例如金门大桥、阿拉伯语等。
* You can breakdown a LLM into a set of ‘features’ which describe different concepts e.g. The Golden Gate Bridge, the arabic language, etc.
* 给定一些文本输入,你可以找出哪些特征在 LLM 的“大脑”中被激活。这发生得快速、廉价且可靠,相同的文本总是激活相同的特征。
* Given some text input, you can figure out which one of these features activate in the LLMs “brain”. This happens quickly, cheaply and reliably, the same text will always activate the same features.
* 在生成时,你可以以特定的强度激活某个特定的特征。
* When generating, you can activate a particular feature at a particular intensity.
* 不同强度下的不同特征会产生不同的效果。例如,阿拉伯语特征在足够强度下会使你的模型说阿拉伯语。将金门大桥特征调到最大会使你的模型认为它就是金门大桥。
* Different features at different intensities will have different effects. For example, the Arabic feature at sufficient intensity will make your model speak Arabic. Turning up the golden gate bridge feature all the way will make your model think it’s the golden gate bridge.
* 你可以根据看到的令牌和其他特征,通过激活和停用特征来实时引导。例如,当你检测到编码序列开始时,你可能会提高语言中与编码相关的某些特征。
* You can steer on the fly by activating and deactivating features based on the tokens and other features you see. For example, when you detect the start of a coding sequence, you might turn up certain features related to coding in the language.
在 Anthropic 黑客马拉松中,Dario Amodei 描述了可解释性可能不是描述他们正在做的事情的正确词汇。
At the Anthropic hackathon, Dario Amoedei described how interpretability is maybe not the right word to describe what they are doing.
可解释性的另一面是引导、精确性和可靠性。如果可解释性实现了其承诺,开发者应该对他们模型有迄今为止不可能达到的控制水平。
The other side of interpretability is steering, preciseness & reliability. If interpretability delivers on its promise, developers should have a level of control over their models that has not been possible thus far.
大语言模型最难解决的问题之一就是捕捉并再现特定的风格。
One of the hardest problems with LLMs is capturing and reproducing a specific style.
在提示词中,你可能会说:“友好且简洁”,但实际上你可能想要的是:“70%友好,50%简洁,80%专业”。
In a prompt you might say: “friendly and concise”, but you might actually want something like: “70% friendly, 50% concise, 80% professional”.
或者你可能想要模仿某个特定的人,却发现提示词“更像[某人]”并不奏效。
Or you may want to emulate a specific person, but find that the prompt: “be more like [person]” doesn’t work.
可解释性和引导能使你将风格分解为一组特征,然后通过激活或停用这些特征来引导模型。
Interpretability and steering allows you to break down a style into a set of features, and then steer the model by activating and deactivating those features.
事实上,研究已经表明你可以引导文本、语音和图像。
In fact, research is already showing that you can steer text, voice and images.
Linus 撰文介绍了通过训练文本嵌入分类器来引导文本生成。
Linus writes about steering text generation by training a text embedding classifier.
Goodfire.ai 提供了一款基于你检测到的特征来引导 llama 模型的工具。
Goodfire.ai provides a tool for steering llama models based on features that you detect.
Hume 能够将语音分解为各个组成部分,然后让您调节诸如“鼻音”或“清脆”之类的特征。
Hume is able to break down a voice into components, and then let you tune up and down features like “nasal” or “crisp”
Gytis 的 Featurelab.xyz 展示了图像如何被分解为不同的特征,以及如何通过组合这些特征来生成图像。
Featurelab.xyz by Gytis shows how images can be broken down into distinct features, and also how you can generate images by composing those features together.
RLHF(及相关技术)迄今为止一直是创建系统个性(例如简洁、有帮助)和施加系统限制(例如不回应与种族主义相关的话题)的主要方式。
RLHF (and related techniques) have so far been the primary way of creating a system personality (e.g. concise, helpful) and imposing system limits (e.g. do not respond about topics related to racism).
问题在于 RLHF 以副作用著称。它可能导致“错误拒绝”,即 AI 拒绝执行它能够做的事情(llama 论文引用 1-5%的新错误拒绝),并且有时会降低质量,例如 RLHF 使其更简洁,也可能导致模型生成的代码注释只是说“填写代码”。
The problem is that RLHF famously has side effects. It can lead to “false refusals” where the AI refuses to do something that it can do (the llama paper cites 1-5% of new false refusals), and sometimes degrades quality, e.g. RLHF to be more concise, can also make a model respond with code comments that just say to fill in the code.
引导可能使我们能够仅在检测到某个特征激活时才做出反应(例如,Anthropic 检测到了“诈骗邮件”特征),然后通过激活另一个特征或停用该特征来实施干预。
Steering potentially allows us to react only if we see a feature activate (eg., Anthropic has detected a ‘scam emails’ feature) and implement an intervention only then by activating another feature, or deactivating that feature.
它还允许我们通过选择特定的个性和响应特征来激活(例如,谄媚特征),从而为模型选择个性。
It also allows us to choose a personality for the model by choosing particular personality and response features to activate (e.g. a sycophantic feature).
甚至已经有了一些通过称为“abliteration”的技术对经过 RLHF 的模型进行“取消审查”的工作。
There has even been some work on ‘uncensoring’ RLHFed models through a technique called abliteration.
最有希望的是,引导发生在推理时,这与 RLHF 不同,这意味着 API 开发者应该能够更具体地选择模型如何被引导,而不是依赖模型提供者在后训练中通过 RLHF 做出影响所有人的大规模笼统选择。
Most promisingly, steering happens at inference time unlike RLHF which means that API devs should be able to choose how the model is steered more specifically, vs having to rely on the model providers making large blanket choices that affect everyone via RLHF in post-training.
在聊天时,用户可能会给你一个_request_,比如“告诉我法国的首都”,或者他们可能表达一个偏好,比如“请更简洁地回复”或“用阿拉伯语说话”。
When chatting may give you a _request_ like “tell me the capital of France” or they may state a preference like “please respond more briefly” or “speak in Arabic”
在长时间的对话过程中,这些偏好可能会在上下文窗口中丢失。但引导允许你识别重要的偏好并永久保存,确保你的 AI 始终更简洁地回复。
Over the course of a long conversation, these preferences may be lost in the context window. But steering allows you to recognize an important preference and save it permanently, making sure that your AI is always responding more briefly.
你可能有一堆你认为属于垃圾邮件的电子邮件示例,但你没有一种简洁的方式向 AI 描述什么是“垃圾邮件”。
You may have a bunch of examples of emails that you would describe as spam, but you don’t have a way of concisely describing to the AI what “spam” means.
目前,你可以给 gpt-4o mini 一组 100 个示例,让它对每个邮件进行分类。这还不错,但有点不可靠且成本较高。
Right now, you could give the gpt-4o mini a set of 100 examples and ask it to classify each email. This is not bad, but it is a bit unreliable and costly.
借助可解释性,你可以给 AI 大量垃圾邮件示例,观察哪些特征被激活得最多,并与非垃圾邮件中激活的特征进行对比。
With interpretability, you can give the AI a large set of examples of spam emails and see what features activate the most, compared to what features activate in non-spam emails.
构建这个“垃圾邮件特征”集将让你本质上创建一个廉价的垃圾邮件分类器,而无需训练一个单独的模型。
Building this set of ‘spam features’ will allow you to essentially create a cheap spam classifier without having to train a seperate model.
操控模型就像脑外科手术,提示则像礼貌地请求。操控模型以过度加权某个特征,不仅会导致其过分强调该特征(例如,Anthropic 描述了在高权重下 Golden Gate Claude 认为自己是金门大桥),而且还会生成不连贯的文本或不遵循语言规则的文本。
Steering is like brain surgery, prompting is like asking politely. Steering a model to overweight a feature can cause it to not just bring it up excessively (e.g. Anthropic describing at high weights how Golden Gate Claude thinks that it is the Golden Gate Bridge), but also to just generate incoherent text or text that doesn’t follow the rules of language.
特征是通过人机混合的方式,通过阅读由这些特征触发的输出列表并找出它们的共同之处来进行标记的。
Features are labelled by a mixture of humans and machines by reading a list of output that is triggered by said features and reading what they might have in common.
然而,这种标记仍处于早期阶段,浏览特征列表时,我偶尔会觉得它们的标记方式与我的理解不一致。
However, this labeling is still early on, browsing lists of features I occasionally find myself disagreeing with how they are labeled.
鉴于可能存在数万或数十万个特征,有些特征很可能会被错误标记或误解,这将使得利用它们进行操控的可靠性降低。
Given that there might tens or hundreds of thousands of features, it seems likely that some will be mislabeled or misunderstood, which will make steering using them less reliable.
有些特征会以复杂的方式激活其他特征(这些有时被称为回路)。我们在激活某个特征时可能会引入其他副作用,就像 RLHF 引入副作用一样。
Some features activate other features in ways that are complex to understand (these are sometimes called circuits). We may be introducing other side effects while activating a feature, in the same way that RLHF introduces side effects.
目前还没有大规模使用特征引导(除非 Anthropic 在幕后进行),因此很难判断它们可能产生什么副作用。
There has been no wide scale use of feature steering (unless Anthropic is doing it under the hood), and so it’s hard to tell what side effects they might produce.
尽管许多可解释性功能仍然有些过早,无法显示出一致的可靠性,但它们为开发者提供了对模型更精细的控制能力。
While a lot of the interpretability features are still a bit too early to show consistent reliability, they promise developers a lot more fine-grained control of models.
作为一种权衡,我们应该预计下一代模型 API 将变得更加强大,但也更加复杂,需要超越简单的提示和检索增强生成(RAG)才能获得我们想要的输出。
As a tradeoff, we should expect the next gen model APIs to get a lot more powerful but also more complicated, requiring more than just prompting and RAG to get the outputs we want.