更新,2026 年 1 月 21 日:我们发布了克劳德宪法的新版本,可在上方按钮找到。语言模型如何决定它愿意回答哪些问题,认为哪些问题不合适?为什么它会鼓励某些行为而劝阻其他行为?语言模型可能具有什么样的“价值观”?这些都是人们努力解决的问题。我们最近发表的关于“宪法 AI”的研究提供了一种答案:通过宪法赋予语言模型明确的价值观,而不是通过大规模人类反馈隐含地确定价值观。这不是一个完美的方法,但它确实使 AI 系统的价值观更容易理解,也更容易根据需要调整。
_Update, Jan 21, 2026:We've published a new version of Claude's constitution, which you can find at the button above._ How does a language model decide which questions it will engage with and which it deems inappropriate? Why will it encourage some actions and discourage others? What “values” might a language model have? These are all questions people grapple with. Our recently published research on “Constitutional AI” provides one answer by giving language models explicit values determined by a constitution, rather than values determined implicitly via large-scale human feedback. This isn’t a perfect approach, but it does make the values of the AI system easier to understand and easier to adjust as needed.
核心贡献 · Key contributions
提出宪法 AI(CAI),使用显式原则而非隐式人类反馈进行模型对齐。 Introduces Constitutional AI (CAI) using explicit principles instead of implicit human feedback for model alignment.
CAI 实现帕累托改进:比 RLHF 既更有帮助又更无害。 CAI achieves Pareto improvement: both more helpful and more harmless than RLHF.
通过使用 AI 反馈训练无害性而无需人类数据,展示了可扩展监督。 Demonstrates scalable oversight by using AI feedback to train harmlessness without human data.
提供透明度:原则易于指定、检查和调整。 Provides transparency: principles are easily specified, inspected, and adjusted.
减少训练中人类接触令人不安内容的需求。 Reduces need for humans to view disturbing content during training.
宪法借鉴联合国宣言、平台指南和安全研究,支持迭代改进。 Constitution draws from UN Declaration, platform guidelines, and safety research, enabling iterative improvement.
局限 · Limitations
宪法反映设计者的选择,可能不代表多元社会价值观。 Constitution reflects designers' choices; may not represent diverse societal values.
CAI 可能产生评判性或烦人的回应,需要额外原则来缓和。 CAI may produce judgmental or annoying responses; requires additional principles to temper.
有效性取决于原则质量;过于具体的原则降低泛化能力。 Effectiveness depends on principle quality; overly specific principles reduce generalization.
宪法 AI 并非万能;关于允许内容的难题依然存在。 Constitutional AI is not a panacea; difficult questions about allowed content remain.
AI 反馈的可扩展性可能面临日益复杂模型输出的挑战。 Scalability of AI feedback may face challenges with increasingly complex model outputs.
论文章节 · Sections(共 7)
Claude 的宪法Claude’s Constitution
背景Context
什么是宪法式 AI?What is Constitutional AI?
宪法包含什么?What's in the Constitution?
这些原则是否有优先级排序?Are these principles prioritized in any way?