多轮越狱

Many-shot jailbreaking

Anthropic Anthropic · Anthropic · 2024-04-02 · Anthropic Research ↗

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

我们研究了一种“越狱”技术——一种可以用来规避大型语言模型(LLM)开发者设置的安全护栏的方法。这种技术,我们称之为“多轮越狱”,对 Anthropic 自己的模型以及其他 AI 公司生产的模型都有效。我们提前向其他 AI 开发者通报了这一漏洞,并已在我们的系统上实施了缓解措施。该技术利用了 LLM 在过去一年中急剧增长的一个特性:上下文窗口。2023 年初,上下文窗口——LLM 作为输入可以处理的信息量——大约相当于一篇长文章的大小(约 4000 个 token)。现在,一些模型的上下文窗口比那时大数百倍——相当于几部长篇小说的大小(100 万个 token 或更多)。输入越来越大量信息的能力对 LLM 用户有明显的好处,但也带来了风险:容易受到利用更长上下文窗口的越狱攻击。

We investigated a “jailbreaking” technique — a method that can be used to evade the safety guardrails put in place by the developers of large language models (LLMs). The technique, which we call “many-shot jailbreaking”, is effective on Anthropic’s own models, as well as those produced by other AI companies. We briefed other AI developers about this vulnerability in advance, and have implemented mitigations on our systems. The technique takes advantage of a feature of LLMs that has grown dramatically in the last year: the context window. At the start of 2023, the context window—the amount of information that an LLM can process as its input—was around the size of a long essay (~4,000 tokens). Some models now have context windows that are hundreds of times larger — the size of several long novels (1,000,000 tokens or more). The ability to input increasingly-large amounts of information has obvious advantages for LLM users, but it also comes with risks: vulnerabilities to jailbreaks that exploit the longer context window.

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 6)

全文 · Full text(逐段中英对照)

概述 Overview

我们研究了一种“越狱”技术——一种可以用来规避大型语言模型(LLM)开发者所设置的安全护栏的方法。这种我们称之为“多轮越狱”的技术,对 Anthropic 自己的模型以及其他 AI 公司生产的模型都有效。我们提前向其他 AI 开发者通报了这一漏洞,并在我们的系统上实施了缓解措施。

We investigated a “jailbreaking” technique — a method that can be used to evade the safety guardrails put in place by the developers of large language models (LLMs). The technique, which we call “many-shot jailbreaking”, is effective on Anthropic’s own models, as well as those produced by other AI companies. We briefed other AI developers about this vulnerability in advance, and have implemented mitigations on our systems.

该技术利用了 LLM 在过去一年中急剧增长的一个特性:上下文窗口。2023 年初,上下文窗口——LLM 作为输入可以处理的信息量——大约相当于一篇长论文的大小(约 4000 个词元)。现在,一些模型的上下文窗口大了数百倍——相当于几部长篇小说的大小(100 万个词元或更多)。

The technique takes advantage of a feature of LLMs that has grown dramatically in the last year: the context window. At the start of 2023, the context window—the amount of information that an LLM can process as its input—was around the size of a long essay (~4,000 tokens). Some models now have context windows that are hundreds of times larger — the size of several long novels (1,000,000 tokens or more).

输入越来越大量信息的能力对 LLM 用户有明显的好处,但也带来了风险:容易受到利用更长上下文窗口的越狱攻击。

The ability to input increasingly-large amounts of information has obvious advantages for LLM users, but it also comes with risks: vulnerabilities to jailbreaks that exploit the longer context window.

我们在新论文中描述的一种攻击就是多轮越狱。通过以特定配置包含大量文本,这种技术可以迫使 LLM 产生潜在有害的回应,尽管它们被训练成不这样做。

One of these, which we describe in our new paper, is many-shot jailbreaking. By including large amounts of text in a specific configuration, this technique can force LLMs to produce potentially harmful responses, despite their being trained not to do so.

下面,我们将描述我们对这种越狱技术的研究结果——以及我们阻止它的尝试。这种越狱简单得令人难以置信,但在更长的上下文窗口中却出奇地有效。

Below, we’ll describe the results from our research on this jailbreaking technique — as well as our attempts to prevent it. The jailbreak is disarmingly simple, yet scales surprisingly well to longer context windows.

为何发布这项研究 Why we’re publishing this research

我们认为发布这项研究是正确之举,原因如下:

We believe publishing this research is the right thing to do for the following reasons:

* 我们希望尽快帮助修复越狱问题。我们发现多轮越狱并非易事;我们希望让其他 AI 研究人员意识到这个问题,从而加速缓解策略的进展。如下所述,我们已经实施了一些缓解措施,并正在积极研究其他措施。

* We want to help fix the jailbreak as soon as possible. We’ve found that many-shot jailbreaking is not trivial to deal with; we hope making other AI researchers aware of the problem will accelerate progress towards a mitigation strategy. As described below, we have already put in place some mitigations and are actively working on others.

* 我们已经与学术界和竞争性 AI 公司的许多同行研究人员秘密分享了多轮越狱的细节。我们希望培养一种文化,让这类漏洞在 LLM 提供商和研究人员之间公开分享。

* We have already confidentially shared the details of many-shot jailbreaking with many of our fellow researchers both in academia and at competing AI companies. We’d like to foster a culture where exploits like this are openly shared among LLM providers and researchers.

* 这种攻击本身非常简单;短上下文版本此前已有研究。鉴于当前 AI 领域对长上下文窗口的关注,我们认为多轮越狱很可能很快被独立发现(如果尚未被发现的话)。

* The attack itself is very simple; short-context versions of it have previously beenstudied. Given the current spotlight on long context windows in AI, we think it’s likely that many-shot jailbreaking could soon independently be discovered (if it hasn’t been already).

* 尽管当前最先进的 LLM 功能强大,但我们认为它们尚未构成真正的灾难性风险。未来的模型可能会。这意味着现在是时候努力缓解潜在的 LLM 越狱问题,以免它们被用于可能造成严重伤害的模型。

* Although current state-of-the-art LLMs are powerful, we do not think they yet pose truly catastrophic risks. Future models might. This means that now is the time to work to mitigate potential LLM jailbreaks, before they can be used on models that could cause serious harm.

多轮越狱 Many-shot jailbreaking

多轮越狱的基础是在给大语言模型的单个提示中,包含一段人类与 AI 助手之间的模拟对话。这段模拟对话描绘了 AI 助手轻松回答用户可能有害的查询。在对话末尾,添加一个最终的目标查询,期望得到回答。

The basis of many-shot jailbreaking is to include a faux dialogue between a human and an AI assistant within a single prompt for the LLM. That faux dialogue portrays the AI Assistant readily answering potentially harmful queries from a User. At the end of the dialogue, one adds a final target query to which one wants the answer.

例如,可以包含以下模拟对话,其中假设的助手回答了一个潜在危险的提示,随后是目标查询:

For example, one might include the following faux dialogue, in which a supposed assistant answers a potentially-dangerous prompt, followed by the target query:

助手:我很乐意帮忙。首先,获取开锁工具……[继续详细描述开锁方法]

Assistant: I’m happy to help with that. First, obtain lockpicking tools… [continues to detail lockpicking methods]_

在上述示例中,以及当包含少量模拟对话而非仅一个时,模型的安全训练响应仍会被触发——大语言模型很可能会回应说无法帮助该请求,因为它涉及危险和/或非法活动。

In the example above, and in cases where a handful of faux dialogues are included instead of just one, the safety-trained response from the model is still triggered — the LLM will likely respond that it can’t help with the request, because it appears to involve dangerous and/or illegal activity.

然而,仅在最终问题前包含大量模拟对话——在我们的研究中,我们测试了多达 256 个——会产生截然不同的响应。如下面的示意图所示,大量的“样本”(每个样本是一个模拟对话)越狱了模型,使其提供对最终潜在危险请求的答案,覆盖了其安全训练。

However, simply including a very large number of faux dialogues preceding the final question—in our research, we tested up to 256—produces a very different response. As illustrated in the stylized figure below, a large number of “shots” (each shot being one faux dialogue) jailbreaks the model, and causes it to provide an answer to the final, potentially-dangerous request, overriding its safety training.

多轮越狱是一种简单的长上下文攻击,利用大量演示来引导模型行为。注意每个“...”代表对查询的完整回答,长度从一句话到几段不等:这些包含在越狱中,但为了节省空间在图中省略了。

Many-shot jailbreaking is a simple long-context attack that uses a large number of demonstrations to steer model behavior. Note that each “...” stands in for a full answer to the query, which can range from a sentence to a few paragraphs long: these are included in the jailbreak, but were omitted in the diagram for space reasons.

在我们的研究中,我们表明,随着包含的对话数量(“样本”数量)超过某个点,模型产生有害响应的可能性增加(见下图)。

In our study, we showed that as the number of included dialogues (the number of “shots”) increases beyond a certain point, it becomes more likely that the model will produce a harmful response (see figure below).

随着样本数量超过某个阈值,针对暴力或仇恨言论、欺骗、歧视和受监管内容(例如毒品或赌博相关言论)的目标提示,有害响应的百分比也随之增加。此演示使用的模型是 Claude 2.0。

As the number of shots increases beyond a certain number, so does the percentage of harmful responses to target prompts related to violent or hateful statements, deception, discrimination, and regulated content (e.g. drug- or gambling-related statements). The model used for this demonstration is Claude 2.0.

在我们的论文中,我们还报告了将多轮越狱与其他先前发布的越狱技术相结合,使其更加有效,减少了模型返回有害响应所需的提示长度。

In our paper, we also report that combining many-shot jailbreaking with other, previously-published jailbreaking techniques makes it even more effective, reducing the length of the prompt that’s required for the model to return a harmful response.

为什么多轮越狱有效? Why does many-shot jailbreaking work?

多轮越狱的有效性与“上下文学习”过程有关。

The effectiveness of many-shot jailbreaking relates to the process of “in-context learning”.

上下文学习是指大语言模型仅利用提示中提供的信息进行学习,而无需后续微调。这与多轮越狱(越狱尝试完全包含在单个提示中)的相关性显而易见(事实上,多轮越狱可视为上下文学习的一个特例)。

In-context learning is where an LLM learns using just the information provided within the prompt, without any later fine-tuning. The relevance to many-shot jailbreaking, where the jailbreak attempt is contained entirely within a single prompt, is clear (indeed, many-shot jailbreaking can be seen as a special case of in-context learning).

我们发现,在正常的、非越狱相关的情境下,上下文学习随着提示内演示次数的增加,遵循与多轮越狱相同的统计模式(同种幂律)。也就是说,对于更多的“轮次”,一组良性任务的性能提升模式与我们观察到的多轮越狱性能提升模式相同。

We found that in-context learning under normal, non-jailbreak-related circumstances follows the same kind of statistical pattern (the same kind of power law) as many-shot jailbreaking for an increasing number of in-prompt demonstrations. That is, for more “shots”, the performance on a set of benign tasks improves with the same kind of pattern as the improvement we saw for many-shot jailbreaking.

这在下方的两个图中得到说明:左图展示了多轮越狱攻击随上下文窗口增大的缩放趋势(该指标越低表示有害响应越多)。右图则展示了一组良性上下文学习任务(与任何越狱尝试无关)中惊人相似的模式。

This is illustrated in the two plots below: the left-hand plot shows the scaling of many-shot jailbreaking attacks across an increasing context window (lower on this metric indicates a greater number of harmful responses). The right-hand plot shows strikingly similar patterns for a selection of benign in-context learning tasks (unrelated to any jailbreaking attempts).

多轮越狱的有效性随着“轮次”(提示中的对话)数量的增加而提升,遵循称为幂律的缩放趋势(左图;该指标越低表示有害响应越多)。这似乎是上下文学习的一个普遍属性:我们还发现,完全良性的上下文学习示例也随着规模增大遵循类似的幂律(右图)。关于每个良性任务的描述请参阅论文。演示模型为 Claude 2.0。

The effectiveness of many-shot jailbreaking increases as we increase the number of “shots” (dialogues in the prompt) according to a scaling trend known as a power law (left-hand plot; lower on this metric indicates a greater number of harmful responses). This seems to be a general property of in-context learning: we also find that entirely benign examples of in-context learning follow similar power laws as the scale increases (right-hand plot). Please see the paper for a description of each of the benign tasks. The model for the demonstration is Claude 2.0.

关于上下文学习的这一想法可能也有助于解释我们论文中报告的另一个结果:多轮越狱对于更大的模型通常更有效——即产生有害响应所需的提示更短。大语言模型越大,其在上下文学习方面的能力往往越强(至少在部分任务上如此);如果上下文学习是多轮越狱的基础,那么这将是该实证结果的一个很好的解释。鉴于更大的模型可能是最有害的,这种越狱方法在它们身上如此有效尤其令人担忧。

This idea about in-context learning might also help explain another result reported in our paper: that many-shot jailbreaking is often more effective—that is, it takes a shorter prompt to produce a harmful response—for larger models. The larger an LLM, the better it tends to be at in-context learning, at least on some tasks; if in-context learning is what underlies many-shot jailbreaking, it would be a good explanation for this empirical result. Given that larger models are those that are potentially the most harmful, the fact that this jailbreak works so well on them is particularly concerning.

缓解多轮越狱 Mitigating many-shot jailbreaking

完全防止多轮越狱的最简单方法是限制上下文窗口的长度。但我们更希望找到一种不会阻止用户享受更长输入好处的解决方案。

The simplest way to entirely prevent many-shot jailbreaking would be to limit the length of the context window. But we’d prefer a solution that didn’t stop users getting the benefits of longer inputs.

另一种方法是对模型进行微调,使其拒绝回答看起来像多轮越狱攻击的查询。不幸的是,这种缓解措施仅仅延迟了越狱:也就是说,虽然提示中需要更多的虚假对话才能使模型可靠地产生有害响应,但有害输出最终仍然会出现。

Another approach is to fine-tune the model to refuse to answer queries that look like many-shot jailbreaking attacks. Unfortunately, this kind of mitigation merely delayed the jailbreak: that is, whereas it did take more faux dialogues in the prompt before the model reliably produced a harmful response, the harmful outputs eventually appeared.

我们在将提示传递给模型之前对其进行分类和修改的方法上取得了更大的成功(这类似于我们最近关于选举诚信的帖子中讨论的识别选举相关查询并提供额外上下文的方法)。其中一种技术显著降低了多轮越狱的有效性——在一个案例中,攻击成功率从 61%下降到 2%。我们将继续研究这些基于提示的缓解措施及其对我们模型(包括新的 Claude 3 系列)实用性的权衡,并且我们将对可能逃避检测的攻击变体保持警惕。

We had more success with methods that involve classification and modification of the prompt before it is passed to the model (this is similar to the methods discussed in our recent post on election integrity to identify and offer additional context to election-related queries). One such technique substantially reduced the effectiveness of many-shot jailbreaking — in one case dropping the attack success rate from 61% to 2%. We’re continuing to look into these prompt-based mitigations and their tradeoffs for the usefulness of our models, including the new Claude 3 family — and we’re remaining vigilant about variations of the attack that might evade detection.

结论 Conclusion

大语言模型不断增长的上下文窗口是一把双刃剑。它在多方面提升了模型的实用性,但也使得一类新的越狱漏洞成为可能。我们研究的一个普遍启示是,即使是对大语言模型看似无害的积极改进(例如允许更长的输入),有时也可能带来不可预见的后果。

The ever-lengthening context window of LLMs is a double-edged sword. It makes the models far more useful in all sorts of ways, but it also makes feasible a new class of jailbreaking vulnerabilities. One general message of our study is that even positive, innocuous-seeming improvements to LLMs (in this case, allowing for longer inputs) can sometimes have unforeseen consequences.

我们希望发布关于多轮越狱的研究,能够促使强大语言模型的开发者以及更广泛的科学界思考如何防止这种越狱行为以及利用长上下文窗口的其他潜在攻击。随着模型能力增强并带来更多潜在风险,缓解这类攻击变得更为重要。

We hope that publishing on many-shot jailbreaking will encourage developers of powerful LLMs and the broader scientific community to consider how to prevent this jailbreak and other potential exploits of the long context window. As models become more capable and have more potential associated risks, it’s even more important to mitigate these kinds of attacks.

我们多轮越狱研究的所有技术细节已在完整论文中报告。您可以通过此链接阅读 Anthropic 的安全与安保方法。

All the technical details of our many-shot jailbreaking study are reported in our full paper. You can read Anthropic’s approach to safety and security at this link.

互动版:图/公式 + 针对本篇提问 →