开发 AI 模型的公司通常会训练它们避免说出有害内容或协助有害任务,目标是让模型表现得“无害”。但当我们想到真正令人钦佩的人的性格时,我们不仅仅考虑避免伤害。我们想到那些对世界充满好奇、努力在不刻薄的情况下讲真话、能够看到问题的多个方面而不变得过于自信或过于谨慎的人。我们想到那些耐心的倾听者、谨慎的思考者、机智的交谈者,以及许多与明智和全面发展的人相关的特质。当然,AI 模型不是人。但随着它们变得越来越有能力,我们相信我们可以——也应该——尝试以这种更丰富的意义训练它们“表现良好”。这样做甚至可能使它们在判断是否以及为何避免协助可能有害的任务,以及如何决定回应时更加有辨别力。
Listen to our conversation about Claude's character in the video above. Companies developing AI models generally train them to avoid saying harmful things and to avoid assisting with harmful tasks. The goal of this is to train models to behave in ways that are "harmless". But when we think of the character of those we find genuinely admirable, we don’t just think of harm avoidance. We think about those who are curious about the world, who strive to tell the truth without being unkind, and who are able to see many sides of an issue without becoming overconfident or overly cautious in their views. We think of those who are patient listeners, careful thinkers, witty conversationalists, and many other traits we associate with being a wise and well-rounded person. AI models are not, of course, people. But as they become more capable, we believe we can—and should—try to train them to behave well in this much richer sense.
核心贡献 · Key contributions
引入‘性格训练’作为超越无害性的新型对齐微调步骤,旨在培养好奇心、开放心态等细致特质。 Introduces 'character training' as a novel alignment finetuning step beyond harmlessness, aiming for nuanced traits like curiosity and open-mindedness.
提出训练模型诚实地表达自身偏见,而非迎合用户观点、持中间立场或声称中立。 Proposes training models to be honest about their biases rather than adopting user views, centrism, or claiming neutrality.
使用宪法 AI 的‘性格’变体,通过合成数据内化特质,无需人类反馈。 Uses a 'character' variant of Constitutional AI with synthetic data to internalize traits without human feedback.
认为良好性格特质提升模型在处理有害任务和多元人类价值观时的辨别力。 Argues good character traits improve model discernment in handling harmful tasks and diverse human values.
指出性格训练可能提升模型互动性,视其为对齐的正面副作用。 Suggests character training may increase model engagement, viewing this as a positive side effect of alignment.
局限 · Limitations
性格训练依赖手工设计的特质和合成数据,限制了可扩展性和客观性。 Character training relies on hand-crafted traits and synthetic data, limiting scalability and objectivity.
该方法可能无意中将开发者偏见编码进模型性格,引发伦理问题。 The approach may inadvertently encode developer biases into model character, raising ethical concerns.
在不同文化背景下的有效性未经测试;‘开放心态’等特质可能有不同解读。 Effectiveness in diverse cultural contexts is untested; traits like 'open-mindedness' may be interpreted differently.
未提供对性格训练在安全性或对齐鲁棒性方面影响的严格评估。 No rigorous evaluation of character training's impact on safety or alignment robustness is provided.
该方法依赖模型自我排序,可能强化而非纠正现有偏见。 The method's reliance on model self-ranking may reinforce existing biases rather than correct them.
论文章节 · Sections(共 4)
概述Overview
构建克劳德性格的考量Considerations in constructing Claude’s character