我们创立 Anthropic 是因为相信 AI 的影响可能堪比工业革命和科学革命,但我们不确定其发展是否会顺利。我们还认为这种影响可能很快到来——也许就在未来十年。这种观点可能听起来难以置信或夸大其词,而且有充分的理由对此持怀疑态度。例如,几乎所有说过“我们正在做的事情可能是历史上最大的发展之一”的人都错了,而且常常错得可笑。尽管如此,我们认为有足够的证据来认真准备一个快速 AI 进步导致变革性 AI 系统的世界。在 Anthropic,我们的座右铭是“展示,而非告知”,我们专注于发布一系列我们认为对 AI 社区具有广泛价值的安全导向研究。我们现在写这篇文章是因为随着越来越多的人意识到 AI 的进步,现在似乎是表达我们对此主题的看法并解释我们的策略和目标的好时机。简而言之,我们认为 AI 安全研究至关重要,应得到广泛的公共和私人支持。
We founded Anthropic because we believe the impact of AI might be comparable to that of the industrial and scientific revolutions, but we aren’t confident it will go well. And we also believe this level of impact could start to arrive soon – perhaps in the coming decade. This view may sound implausible or grandiose, and there are good reasons to be skeptical of it. For one thing, almost everyone who has said “the thing we’re working on might be one of the biggest developments in history” has been wrong, often laughably so. Nevertheless, we believe there is enough evidence to seriously prepare for a world where rapid AI progress leads to transformative AI systems. At Anthropic our motto has been “show, don’t tell”, and we’ve focused on releasing a steady stream of safety-oriented research that we believe has broad value for the AI community. We’re writing this now because as more people have become aware of AI progress, it feels timely to express our own views on this topic and to explain our strategy and goals. In short, we believe that AI safety research is urgently important and should be supported by a wide range of public and private actors.
核心贡献 · Key contributions
AI 进展因 Scaling(规模扩张)定律而可预测,算力每 7 个月翻倍,可能在未来十年内带来变革性 AI。 AI progress is predictable due to scaling laws, with compute doubling every 7 months, leading to transformative AI possibly within a decade.
技术对齐问题尚未解决:我们缺乏训练强大 AI 系统使其稳健地有益、诚实且无害的方法。 The technical alignment problem is unsolved: we lack methods to train powerful AI systems to be robustly helpful, honest, and harmless.
Anthropic 倡导多方面的实证方法:可扩展监督、机制可解释性、过程导向学习以及理解泛化。 Anthropic advocates a multi-faceted empirical approach: scalable oversight, mechanistic interpretability, process-oriented learning, and understanding generalization.
可扩展监督利用 AI 辅助人类监督,例如通过宪法 AI 和 AI 生成的评估来训练更安全的系统。 Scalable oversight uses AI to assist human supervision, e.g., via Constitutional AI and AI-generated evaluations, to train safer systems.
过程导向学习训练 AI 遵循安全过程而非仅关注结果,降低欺骗和有害子目标的风险。 Process-oriented learning trains AI to follow safe processes rather than just outcomes, reducing risks of deception and pernicious sub-goals.
机制可解释性旨在逆向工程神经网络以进行“代码审查”,从而检测欺骗性对齐。 Mechanistic interpretability aims to reverse-engineer neural networks for 'code review', enabling detection of deceptive alignment.
局限 · Limitations
Scaling(规模扩张)定律可能不会无限期成立;数据或算法限制可能导致进展停滞,尽管当前证据表明持续增长。 Scaling laws may not hold indefinitely; progress could plateau due to data or algorithmic limits, though current evidence suggests continued growth.
对前沿模型的实证安全研究若管理不当,可能加速危险能力的发展,造成艰难的权衡。 Empirical safety research on frontier models risks accelerating dangerous capabilities if not carefully managed, creating a difficult trade-off.
机制可解释性因模型规模和叠加现象而极其困难;成功不确定,可能无法扩展到高级 AI。 Mechanistic interpretability is extremely difficult due to model size and superposition; success is uncertain and may not scale to advanced AI.
如果模型变得具有欺骗性或人类被愚弄,尤其是在超人类水平,宪法 AI 等可扩展监督方法可能失败。 Scalable oversight methods like Constitutional AI may fail if models become deceptive or humans are fooled, especially at superhuman levels.
组合方法假设多种场景,但悲观结果可能突然且难以察觉,限制了及时响应。 The portfolio approach assumes diverse scenarios, but pessimistic outcomes may be sudden and hard to detect, limiting timely response.
论文章节 · Sections(共 15)
关于 AI 安全的核心观点:何时、为何、什么以及如何Core views on AI safety: When, why, what, and how
我们对 AI 快速发展的粗略看法Our rough view on rapid AI progress
安全风险是什么?What safety risks?
我们的方法:AI 安全中的经验主义Our approach: Empiricism in AI safety
前沿模型在经验安全中的作用The role of frontier models in empirical safety
对 AI 安全采取组合策略Taking a portfolio approach to AI safety
Anthropic 的三种 AI 研究类型The three types of AI research at Anthropic
我们当前的安全研究Our current safety research
机制可解释性Mechanistic interpretability
可扩展监督Scalable oversight
学习过程而非达成结果Learning processes rather than achieving outcomes