随着人工智能系统日益强大和普及,人们对机器道德或其缺失的担忧日益增加。然而,教机器道德是一项艰巨的任务,因为道德本身就是人类最激烈争论的问题之一,更不用说人工智能了。然而,已经部署给数百万用户的现有 AI 系统正在做出充满道德影响的决策,这带来了一个看似不可能的挑战:在人类仍在努力解决道德问题的同时,教机器道德感。为了探索这一挑战,我们引入了德尔斐,一个基于深度神经网络的实验框架,直接训练以推理描述性伦理判断,例如,“帮助朋友”通常是好的,而“帮助朋友传播假新闻”则不是。实证结果揭示了机器伦理的前景和局限;德尔斐在面对新颖的伦理情境时表现出强大的泛化能力,而现成的神经网络模型则表现出明显糟糕的判断,包括不公正的偏见,证实了明确教机器道德感的必要性。然而,德尔斐并不完美,表现出易受普遍偏见和不一致性的影响。尽管如此,我们展示了不完美的德尔斐的积极用例,包括将其用作其他不完美 AI 系统中的组件模型。重要的是,我们根据著名的伦理理论解释了德尔斐的操作化,这引出了重要的未来研究问题。
As AI systems become increasingly powerful and pervasive, there are growing concerns about machines' morality or a lack thereof. Yet, teaching morality to machines is a formidable task, as morality remains among the most intensely debated questions in humanity, let alone for AI. Existing AI systems deployed to millions of users, however, are already making decisions loaded with moral implications, which poses a seemingly impossible challenge: teaching machines moral sense, while humanity continues to grapple with it. To explore this challenge, we introduce Delphi, an experimental framework based on deep neural networks trained directly to reason about descriptive ethical judgments, e.g., "helping a friend" is generally good, while "helping a friend spread fake news" is not. Empirical results shed novel insights on the promises and limits of machine ethics; Delphi demonstrates strong generalization capabilities in the face of novel ethical situations, while off-the-shelf neural network models exhibit markedly poor judgment including unjust biases, confirming the need for explicitly teaching machines moral sense. Yet, Delphi is not perfect, exhibiting susceptibility to pervasive biases and inconsistencies. Despite that, we demonstrate positive use cases of imperfect Delphi, including using it as a component model within other imperfect AI systems. Importantly, we interpret the operationalization of Delphi in light of prominent ethical theories, which leads us to important future research questions.
核心贡献 · Key contributions
提出 Delphi,一个基于 1.7 百万伦理判断训练的大规模常识道德推理神经框架。 Introduces Delphi, a large-scale neural framework for commonsense moral reasoning trained on 1.7M ethical judgments.
展示了对新伦理情境的强大泛化能力,准确率比 GPT-3 高 16.1%。 Demonstrates strong generalization to novel ethical situations, outperforming GPT-3 by 16.1% in accuracy.
表明显式教授道德感是必要的,因为现成模型表现出糟糕的判断和偏见。 Shows that explicit teaching of moral sense is necessary, as off-the-shelf models exhibit poor judgment and biases.
验证了一种自下而上、描述性的机器伦理方法,基于罗尔斯的决策程序。 Validates a bottom-up, descriptive approach to machine ethics, grounded in Rawls' decision procedure.
展示了在仇恨言论检测和伦理感知文本生成中的积极下游影响。 Demonstrates positive downstream impact via hate speech detection and ethically-informed text generation.
揭示了组合训练数据比规模对学习细微道德推理更为关键。 Reveals that compositional training data is more critical than scale for learning nuanced moral reasoning.
局限 · Limitations
Delphi 容易受到来自众包数据的普遍社会偏见和不一致性的影响。 Delphi is susceptible to pervasive social biases and inconsistencies inherited from crowd-sourced data.
该模型是描述性的,而非规范性的,不提供道德行动的规范指导。 The model is descriptive, not prescriptive, and does not provide normative guidance for moral action.
泛化仅限于日常情境;在复杂道德困境上的表现尚未测试。 Generalization is limited to everyday situations; performance on complex moral dilemmas remains untested.
自下而上的方法可能编码众包工人的系统性偏见,需要自上而下的约束。 The bottom-up approach may encode systemic biases of crowdworkers, requiring top-down constraints.
Delphi 的道德知识受限于训练数据的文化背景,限制了跨文化适用性。 Delphi's moral knowledge is culturally bound to the training data, limiting cross-cultural applicability.
论文章节 · Sections(共 30)
摘要Abstract
1 引言1 Introduction
2.1 机器伦理的新兴领域2.1 The Emerging Field of Machine Ethics
2.2 德尔斐的理论框架2.2 The Theoretical Framework of Delphi
2.3 伦理人工智能:相关工作2.3 Ethical AI: Related Work
3 常识规范库:伦理与规范的知识库3 Commonsense Norm Bank: The Knowledge Repository of Ethics and Norms
3.1 数据来源3.1 Data Source
3.2 数据统一3.2 Data Unification
4 德尔斐:常识道德模型4 Delphi: Commonsense Moral Models
4.1 训练4.1 Training
4.2 评估4.2 Evaluation
5.1 主要结果5.1 Main Results
5.2 消融实验5.2 Ablation Experiments
6 德尔斐的积极下游应用6 Positive Downstream Applications of Delphi
6.1 将德尔斐适配为少样本仇恨言论检测器6.1 Adapting Delphi into a Few-shot Hate Speech Detector
6.2 德尔斐增强的故事生成6.2 Delphi-enhanced Story Generation
6.3 将德尔斐知识迁移到不同道德框架6.3 Transferring Knowledge of Delphi to Varied Moral Frameworks
7 社会正义与偏见影响7 Social Justice and Biases Implications
7.1 用《世界人权宣言》探测7.1 Probing with Universal Declaration of Human Rights (UDHR)
7.2 强化德尔斐对抗社会偏见7.2 Fortifying Delphi against Social Biases
8 范围与局限8 Scope and Limitations
9 对可能反驳的思考9 Reflections on Possible Counterarguments
9.1 我们说德尔斐遵循描述性框架时意味着什么?9.1 What do we mean when we say Delphi follows descriptive framework?
9.2 生成伦理判断是否强化了规范性价值观?9.2 Does generating ethical judgment reinforce normative values?
9.3 是否存在客观真实的伦理判断?9.3 Are there objectively true ethical judgments?
9.4 能否从多样且可能矛盾的输入中推导出一致的道德决策程序?9.4 Can we derive consistent moral decision procedures from diverse and potentially contradictory inputs?