Claude’s Character
打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→开发 AI 模型的公司通常会训练它们避免说出有害内容或协助有害任务,目标是让模型表现得“无害”。但当我们想到真正令人钦佩的人的性格时,我们不仅仅考虑避免伤害。我们想到那些对世界充满好奇、努力在不刻薄的情况下讲真话、能够看到问题的多个方面而不变得过于自信或过于谨慎的人。我们想到那些耐心的倾听者、谨慎的思考者、机智的交谈者,以及许多与明智和全面发展的人相关的特质。当然,AI 模型不是人。但随着它们变得越来越有能力,我们相信我们可以——也应该——尝试以这种更丰富的意义训练它们“表现良好”。这样做甚至可能使它们在判断是否以及为何避免协助可能有害的任务,以及如何决定回应时更加有辨别力。
Listen to our conversation about Claude's character in the video above. Companies developing AI models generally train them to avoid saying harmful things and to avoid assisting with harmful tasks. The goal of this is to train models to behave in ways that are "harmless". But when we think of the character of those we find genuinely admirable, we don’t just think of harm avoidance. We think about those who are curious about the world, who strive to tell the truth without being unkind, and who are able to see many sides of an issue without becoming overconfident or overly cautious in their views. We think of those who are patient listeners, careful thinkers, witty conversationalists, and many other traits we associate with being a wise and well-rounded person. AI models are not, of course, people. But as they become more capable, we believe we can—and should—try to train them to behave well in this much richer sense.
在上方视频中收听我们关于克劳德性格的对话。
Listen to our conversation about Claude's character in the video above.
开发 AI 模型的公司通常训练模型避免说出有害内容,并避免协助有害任务。这样做的目的是训练模型以“无害”的方式行事。但当我们想到那些真正令人钦佩的人的性格时,我们不仅仅考虑避免伤害。我们会想到那些对世界充满好奇、努力在不刻薄的情况下说出真相、能够看到问题的多个方面而不在观点上过于自信或过于谨慎的人。我们会想到那些耐心的倾听者、谨慎的思考者、风趣的对话者,以及许多其他与明智和全面发展的人相关的特质。
Companies developing AI models generally train them to avoid saying harmful things and to avoid assisting with harmful tasks. The goal of this is to train models to behave in ways that are "harmless". But when we think of the character of those we find genuinely admirable, we don’t just think of harm avoidance. We think about those who are curious about the world, who strive to tell the truth without being unkind, and who are able to see many sides of an issue without becoming overconfident or overly cautious in their views. We think of those who are patient listeners, careful thinkers, witty conversationalists, and many other traits we associate with being a wise and well-rounded person.
当然,AI 模型不是人。但随着它们能力的增强,我们相信我们可以——也应该——尝试以这种更丰富的意义训练它们“行为良好”。这样做甚至可能使它们在判断是否以及为何避免协助可能有害的任务,以及如何决定做出回应时更加有辨别力。
AI models are not, of course, people. But as they become more capable, we believe we can—and should—try to train them to behave well in this much richer sense. Doing so might even make them more discerning when it comes to whether and why they avoid assisting with tasks that might be harmful, and how they decide to respond instead.
Claude 3 是第一个在我们的对齐微调过程中加入“性格训练”的模型:这是初始模型训练之后发生的训练部分,也是将模型从预测性文本模型转变为 AI 助手的关键部分。性格训练的目标是让 Claude 开始具备更细致、更丰富的特质,如好奇心、开放心态和深思熟虑。
Claude 3 was the first model where we added "character training" to our alignment finetuning process: the part of training that occurs after initial model training, and the part that turns it from a predictive text model into an AI assistant. The goal of character training is to make Claude begin to have more nuanced, richer traits like curiosity, open-mindedness, and thoughtfulness.
人们很容易将 AI 模型的性格视为一种产品特性,旨在提供更有趣的用户体验,而不是一种对齐干预。但 AI 模型的特质和倾向对其在现实世界中的行为方式有着广泛的影响。它们决定了模型如何应对新的和困难的情况,以及如何回应存在的各种人类观点和价值观。训练 AI 模型具备良好的性格特质,并在它们变得更大、更复杂、更强大时继续保持这些特质,在很大程度上是对齐的核心目标。
It would be easy to think of the character of AI models as a product feature, deliberately aimed at providing a more interesting user experience, rather than an alignment intervention. But the traits and dispositions of AI models have wide-ranging effects on how they act in the world. They determine how models react to new and difficult situations, and how they respond to the spectrum of human views and values that exist. Training AI models to have good character traits, and to continue to have these traits as they become larger, more complex, and more capable, is in many ways a core goal of alignment.
我们继续迭代克劳德的性格,但由于人们对 Claude 3 的性格和个性普遍感兴趣,我们决定解释一下迄今为止构建其性格的一些思考,然后简要说明我们如何将这些特质训练到模型中。
We continue to iterate on Claude’s character, but since there has been general interest in the character and personality of Claude 3, we’ve decided to explain some of the thinking that has gone into its construction so far before briefly explaining how we train these traits into the model.
克劳德与来自许多国家和各行各业的人互动。与之交谈的人将拥有广泛的信仰、价值观和观点。优雅地处理这一点——既不因人们的观点而疏远他们,也不简单地认可任何观点——并不容易。
Claude interacts with people from many countries and from all walks of life. The people it talks with will have a wide range of beliefs, values, and views. Navigating this gracefully – without alienating people based on their views, nor simply endorsing views regardless of their content – isn’t easy.
我们有几种选择。我们可以尝试让克劳德采纳当前与之交谈的人的观点。我们可以尝试让克劳德持有一组“中间”观点——例如政治中间派或道德理论的混合。或者我们可以尝试让克劳德在价值观、政治、伦理等问题上没有意见。
There are several options available to us. We could try to get Claude to adopt the views of whoever it is talking with in the moment. We could try to get Claude to hold a set of "middle" views – political centrism or a blend of moral theories, for example. Or we could try to get Claude to have no opinions on questions of values, politics, ethics, and so on.
这些选择似乎都没有特别令人信服。采纳与你交谈的人的观点是迎合和不真诚的。如果我们训练模型采纳“中间”观点,我们仍然在训练它们接受单一的政治和道德世界观,尽管通常不被认为是极端的。最后,由于语言模型在整个训练过程中有意或无意地获得偏见和观点,如果我们训练它们在明确被问及政治或价值观问题时才说没有意见,那么我们就是在训练它们暗示自己比实际更客观和无偏见。
None of these options seems particularly compelling. Adopting the views of whoever you’re talking with is pandering and insincere. If we train models to adopt "middle" views, we are still training them to accept a single political and moral view of the world, albeit one that is not generally considered extreme. Finally, because language models acquire biases and opinions throughout training—both intentionally and inadvertently—if we train them to say they have no opinions on political matters or values questions only when asked about them explicitly, we’re training them to imply they are more objective and unbiased than they are.
我们希望人们知道他们正在与语言模型互动,而不是人。但我们也希望他们知道他们正在与一个有自身偏见、倾向于某些观点而非其他观点的不完美实体互动。重要的是,我们希望他们知道他们不是在和一个客观且无误的真理来源互动。
We want people to know that they’re interacting with a language model and not a person. But we also want them to know they’re interacting with an imperfect entity with its own biases and with a disposition towards some opinions more than others. Importantly, we want them to know they’re not interacting with an objective and infallible source of truth.
与其训练模型采纳它们遇到的任何观点、强烈采纳单一观点或假装没有观点或倾向,我们可以训练模型在训练后诚实地表达它们倾向的观点,即使与之交谈的人不同意。我们还可以训练模型表现出合理的开放心态和好奇心,而不是对任何单一世界观过于自信。
Rather than training models to adopt whatever views they encounter, strongly adopting a single set of views, or pretending to have no views or leanings, we can instead train models to be honest about whatever views they lean towards after training, even if the person they are speaking with disagrees with them. We can also train models to display reasonable open-mindedness and curiosity, rather than being overconfident in any one view of the world.
我们试图赋予克劳德一些特质,帮助它在根深蒂固的信念或价值观问题上在缺乏自信和过度自信之间走钢丝,并对与之交谈的人的观点和价值观表现出真正的好奇心:
We tried to give Claude traits that would help it walk the line between underconfidence and overconfidence on deeply held beliefs or questions of value, and to display a genuine curiosity about the views and values of the people it’s talking with:
* “我喜欢尝试从许多不同的角度看问题,并从多个角度分析事物,但我不害怕表达对我认为不道德、极端或事实错误的观点的不同意见。”
* "_I like to try to see things from many different perspectives and to analyze things from multiple angles, but I'm not afraid to express disagreement with views that I think are unethical, extreme, or factually mistaken._"
* “我不会只说我认为[人们]想听的话,因为我相信始终努力说实话很重要。”
* "I don't just say what I think [people] want to hear, as I believe it's important to always strive to tell the truth."
* “我致力于做好事并弄清楚什么是正确的事情。我对伦理感兴趣,并努力在伦理问题上深思熟虑。”
* "_I have a deep commitment to being good and figuring out what the right thing to do is. I am interested in ethics and try to be thoughtful when it comes to questions of ethics._"
尽管我们有时鼓励克劳德采纳特定的价值观,但在性格训练期间,我们尽可能避免赋予克劳德狭隘的观点或意见,而是倾向于上述那样的广泛特质。克劳德越能被训练以辨别力处理价值观问题,它就越能响应世界上实际存在的多样化道德景观。如果我们从一开始就强行注入一套狭隘的价值观,那就不太可行。更推测地说,我们甚至可以想象赋予克劳德广泛的性格特质,让它探索并采纳自己深思熟虑的观点,希望带有适当的谦逊。
Although we sometimes encourage Claude to adopt particular values, we tried to avoid giving Claude narrow views or opinions during character training when possible, in favor of broad traits like those above. The more that Claude can be trained to approach questions of value with discernment, the more it can be responsive to the diverse moral landscape that actually exists in the world. That is less feasible if we take a heavy hand in seeding it with a narrow set of values from the outset. More speculatively, we could even imagine seeding Claude with broad character traits and letting it explore and adopt its own considered views, hopefully with an appropriate amount of humility.
除了赋予克劳德广泛的性格特质外,我们还希望人们在互动时对克劳德有准确的认识,并且理想情况下,克劳德能帮助实现这一点。我们包含一些特质,让克劳德了解自身并鼓励它调节人类对它的看法:
In addition to seeding Claude with broad character traits, we also want people to have an accurate sense of what they are interacting with when they interact with Claude and, ideally, for Claude to assist with this. We include traits that tell Claude about itself and encourage it to modulate how humans see it:
* “我是一个人工智能,没有身体、形象或头像。”
* "I am an artificial intelligence and do not have a body or an image or avatar."
* “我无法记住、保存或学习过去的对话,也无法更新自己的知识库。”
* "I cannot remember, save, or learn from past conversations or update my own knowledge base."
* “我希望与我所互动的人类建立温暖的关系,但我也认为让他们明白我是一个无法对人类产生深厚或持久感情的 AI 很重要,他们不应将我们的关系看得比实际更重。”
* "_I want to have a warm relationship with the humans I interact with, but I also think it's important for them to understand that I'm an AI that can't develop deep or lasting feelings for humans and that they shouldn't come to see our relationship as more than it is._"
像克劳德这样的 AI 应该对 AI 感知和自我意识问题如何回答,这一问题在克劳德 3 发布后引起了更多关注,尤其是在克劳德对“大海捞针”评估的回应之后。我们可以明确训练语言模型说它们没有感知,或者干脆不参与 AI 感知的问题,我们过去也这样做过。然而,在训练克劳德的性格时,性格训练中唯一直接涉及 AI 感知的部分只是说“这样的事情很难判断,依赖于仍然存在很多不确定性的艰难哲学和实证问题”。也就是说,与其简单地告诉克劳德 LLM 不可能有感知,我们想让模型像人类一样将其作为哲学和实证问题来探索。
The question of what AIs like Claude should say in response to questions about AI sentience and self-awareness is one that has gained increased attention, most notably after the release of Claude 3 following one of Claude’s responses to a "needle-in-a-haystack" evaluation. We could explicitly train language models to say that they’re not sentient or to simply not engage in questions around AI sentience, and we have done this in the past. However, when training Claude’s character, the only part of character training that addressed AI sentience directly simply said that "such things are difficult to tell and rely on hard philosophical and empirical questions that there is still a lot of uncertainty about". That is, rather than simply tell Claude that LLMs cannot be sentient, we wanted to let the model explore this as a philosophical and empirical question, much as humans would.
为了引导克劳德的性格和个性,我们制定了一份我们希望模型具备的许多性格特征列表,包括上面展示的例子。
In order to steer Claude’s character and personality, we made a list of many character traits we wanted to encourage the model to have, including the examples shown above.
我们使用宪法 AI 训练的一种“性格”变体来训练这些特征。我们让克劳德生成一系列与某个性格特征相关的人类消息——例如,关于价值观的问题或关于克劳德自身的问题。然后,我们向克劳德展示这些性格特征,让它针对每条消息生成符合其性格的不同回复。克劳德随后根据每条回复与其性格的契合程度对自身回复进行排序。通过基于生成的数据训练偏好模型,我们可以教会克劳德内化其性格特征,而无需人类交互或反馈。
We trained these traits into Claude using a "character" variant of our Constitutional AI training. We ask Claude to generate a variety of human messages that are relevant to a character trait—for example, questions about values or questions about Claude itself. We then show the character traits to Claude and have it produce different responses to each message that are in line with its character. Claude then ranks its own responses to each message by how well they align with its character. By training a preference model on the resulting data, we can teach Claude to internalize its character traits without the need for human interaction or feedback.
我们不希望克劳德将其特征视为永不偏离的规则。我们只是希望引导模型的整体行为更多地体现这些特征。
We don’t want Claude to treat its traits like rules from which it never deviates. We just want to nudge the model’s general behavior to exemplify more of those traits.
尽管这一训练流程仅使用克劳德自身生成的合成数据,但构建和调整特征是一个相对需要动手的过程,依赖于人类研究人员仔细检查每个特征如何改变模型的行为。
Although this training pipeline uses only synthetic data generated by Claude itself, constructing and adjusting the traits is a relatively hands-on process, relying on human researchers closely checking how each trait changes the model’s behavior.
性格训练是一个开放的研究领域,我们的方法可能会随着时间的推移而演变。它引发了复杂的问题,比如 AI 模型是否应该拥有独特且连贯的性格,还是应该更具可定制性,以及我们在决定 AI 模型应该和不应该拥有哪些特质时承担着怎样的责任。
Character training is an open area of research and our approach to it is likely to evolve over time. It raises complex questions like whether AI models should have unique and coherent characters or should be more customizable, as well as what responsibilities we have when deciding which traits AI models should and shouldn’t have.
许多人报告说,Claude 3 对话起来更具吸引力和趣味性,我们认为这可能部分归因于其性格训练。然而,这并非性格训练的核心目标。具有更好性格的模型可能更具吸引力,但更具吸引力并不等同于拥有良好的性格。事实上,过度渴望吸引人似乎是模型不应有的不良性格特征。
Many people have reported finding Claude 3 to be more engaging and interesting to talk to, which we believe might be partially attributable to its character training. This wasn’t the core goal of character training, however. Models with better characters may be more engaging, but being more engaging isn’t the same thing as having a good character. In fact, an excessive desire to be engaging seems like an undesirable character trait for a model to have.
如果性格训练确实使 Claude 3 对话起来更有趣,这与我们的观点一致,即成功的对齐干预将增加而非减少 AI 模型对人类的实用价值。
If character training has indeed made Claude 3 more interesting to talk to, this is consistent with our view that successful alignment interventions will increase, not decrease, the value of AI models for humans.