克劳德的宪法 / Anthropic

Claude’s Constitution \ Anthropic

Anthropic Anthropic · Anthropic · 2023-05-09 · Anthropic ↗

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

更新,2026 年 1 月 21 日:我们发布了克劳德宪法的新版本,可在上方按钮找到。语言模型如何决定它愿意回答哪些问题,认为哪些问题不合适?为什么它会鼓励某些行为而劝阻其他行为?语言模型可能具有什么样的“价值观”?这些都是人们努力解决的问题。我们最近发表的关于“宪法 AI”的研究提供了一种答案:通过宪法赋予语言模型明确的价值观,而不是通过大规模人类反馈隐含地确定价值观。这不是一个完美的方法,但它确实使 AI 系统的价值观更容易理解,也更容易根据需要调整。

_Update, Jan 21, 2026:We've published a new version of Claude's constitution, which you can find at the button above._ How does a language model decide which questions it will engage with and which it deems inappropriate? Why will it encourage some actions and discourage others? What “values” might a language model have? These are all questions people grapple with. Our recently published research on “Constitutional AI” provides one answer by giving language models explicit values determined by a constitution, rather than values determined implicitly via large-scale human feedback. This isn’t a perfect approach, but it does make the values of the AI system easier to understand and easier to adjust as needed.

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 7)

阅读逐段中英对照全文 →