叠加的玩具模型

Toy Models of Superposition

Anthropic Anthropic · · 2022-09-21 · arXiv:2209.10652 ↗ · 被引 902

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

神经网络通常将许多不相关的概念打包到单个神经元中——这是一个被称为“多义性”的令人困惑的现象,使得可解释性更具挑战性。本文提供了一个玩具模型,在该模型中多义性可以被完全理解,它是模型在“叠加”中存储额外稀疏特征的结果。我们证明了相变的存在、与均匀多面体几何的惊人联系,以及与对抗性样本关联的证据。我们还讨论了对机制可解释性的潜在影响。

Neural networks often pack many unrelated concepts into a single neuron - a puzzling phenomenon known as 'polysemanticity' which makes interpretability much more challenging. This paper provides a toy model where polysemanticity can be fully understood, arising as a result of models storing additional sparse features in "superposition." We demonstrate the existence of a phase change, a surprising connection to the geometry of uniform polytopes, and evidence of a link to adversarial examples. We also discuss potential implications for mechanistic interpretability.

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 24)

阅读逐段中英对照全文 →