论证可解释性对 AI 安全至关重要,能够检测欺骗、权力寻求和越狱行为。 Argues interpretability is crucial for AI safety, enabling detection of deception, power-seeking, and jailbreaks.
描述机械可解释性的进展:特征、稀疏自编码器和用于追踪模型推理的电路。 Describes progress in mechanistic interpretability: features, sparse autoencoders, and circuits for tracing model reasoning.
展示实际应用:红队/蓝队实验表明可解释性工具能发现对齐缺陷。 Demonstrates practical use: red-team/blue-team experiments show interpretability tools can find alignment flaws.
提出可解释性作为对齐的独立测试集,补充 RLHF 等训练技术。 Proposes interpretability as an independent test set for alignment, complementing training techniques like RLHF.
预测可解释性可能在 5-10 年内达到可靠的“AI MRI”,但 AI 进展可能更快。 Predicts interpretability could reach reliable 'MRI for AI' within 5-10 years, but AI advances may outpace it.
建议政策行动:透明度立法、出口管制和增加研究投入以争取时间。 Recommends policy actions: transparency legislation, export controls, and increased research investment to buy time.
局限 · Limitations
可解释性方法目前扩展性差;仅识别了小模型中的一小部分特征。 Interpretability methods currently scale poorly; only a fraction of features in small models are identified.
叠加使特征纠缠;稀疏自编码器仅能部分解开。 Superposition makes features tangled; sparse autoencoders only partially disentangle them.
电路发现仍手动且有限;自动化发现数百万电路是开放挑战。 Circuit discovery remains manual and limited; automating it for millions of circuits is an open challenge.
可解释性可能无法检测所有风险,尤其是当模型主动隐藏恶意意图时。 Interpretability may not detect all risks, especially if models actively hide malicious intent.
政策建议依赖地缘政治假设;出口管制可能无效或不可持续。 Policy recommendations rely on geopolitical assumptions; export controls may not be effective or sustainable.
论文章节 · Sections(共 6)
目录Contents
可解释性的紧迫性The Urgency of Interpretability
无知之险The Dangers of Ignorance
机械可解释性简史A Brief History of Mechanistic Interpretability