克劳德的性格

Claude’s Character

Anthropic Anthropic · Anthropic · 2024-06-08 · Anthropic Research ↗

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

开发 AI 模型的公司通常会训练它们避免说出有害内容或协助有害任务,目标是让模型表现得“无害”。但当我们想到真正令人钦佩的人的性格时,我们不仅仅考虑避免伤害。我们想到那些对世界充满好奇、努力在不刻薄的情况下讲真话、能够看到问题的多个方面而不变得过于自信或过于谨慎的人。我们想到那些耐心的倾听者、谨慎的思考者、机智的交谈者,以及许多与明智和全面发展的人相关的特质。当然,AI 模型不是人。但随着它们变得越来越有能力,我们相信我们可以——也应该——尝试以这种更丰富的意义训练它们“表现良好”。这样做甚至可能使它们在判断是否以及为何避免协助可能有害的任务,以及如何决定回应时更加有辨别力。

Listen to our conversation about Claude's character in the video above. Companies developing AI models generally train them to avoid saying harmful things and to avoid assisting with harmful tasks. The goal of this is to train models to behave in ways that are "harmless". But when we think of the character of those we find genuinely admirable, we don’t just think of harm avoidance. We think about those who are curious about the world, who strive to tell the truth without being unkind, and who are able to see many sides of an issue without becoming overconfident or overly cautious in their views. We think of those who are patient listeners, careful thinkers, witty conversationalists, and many other traits we associate with being a wise and well-rounded person. AI models are not, of course, people. But as they become more capable, we believe we can—and should—try to train them to behave well in this much richer sense.

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 4)

阅读逐段中英对照全文 →