映射大语言模型的思维

Mapping the mind of a large language model

Anthropic Anthropic · Anthropic · 2024-05-21 · Anthropic Research ↗

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

今天我们报告了在理解 AI 模型内部工作原理方面的一项重大进展。我们确定了数百万个概念是如何在我们部署的大语言模型之一 Claude Sonnet 内部表示的。这是首次对现代、生产级大语言模型进行详细的内窥。这一可解释性发现未来可能有助于使 AI 模型更安全。我们通常将 AI 模型视为黑箱:输入内容,输出响应,但不清楚模型为何给出特定响应而非其他。这使得我们难以信任这些模型的安全性:如果我们不知道它们如何工作,如何知道它们不会给出有害、有偏见、不真实或其他危险的响应?我们如何信任它们会是安全可靠的?打开黑箱并不一定有帮助:模型的内部状态——模型在写出响应前“思考”的内容——由一长串数字(“神经元激活值”)组成,没有明确含义。通过与像 Claude 这样的模型交互,很明显它能够理解并运用广泛的概念——但我们无法通过直接观察神经元来辨别它们。

_Today we report a significant advance in understanding the inner workings of AI models. We have identified how millions of concepts are represented inside Claude Sonnet, one of our deployed large language models. This is the first ever detailed look inside a modern, production-grade large language model._ This interpretability discovery could, in future, help us make AI models safer. We mostly treat AI models as a black box: something goes in and a response comes out, and it's not clear why the model gave that particular response instead of another.

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 1)

阅读逐段中英对照全文 →