人工智能安全中的具体问题

Concrete Problems in AI Safety

达里奥·阿莫迪 Dario Amodei · Google · 2016-06-21 · arXiv:1606.06565 ↗ · 被引 3244

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

机器学习和人工智能的快速发展使人们越来越关注 AI 技术对社会的潜在影响。本文讨论了一种潜在影响:机器学习系统中的事故问题,即由于现实世界 AI 系统设计不当而可能出现的意外和有害行为。我们列出了与事故风险相关的五个实际研究问题,根据问题源于目标函数错误(“避免副作用”和“避免奖励黑客”)、目标函数评估过于昂贵(“可扩展监督”),还是学习过程中的不良行为(“安全探索”和“分布偏移”)进行分类。我们回顾了这些领域的先前工作,并提出了与前沿 AI 系统相关的研究方向。最后,我们考虑了如何最有效地思考前瞻性 AI 应用的安全性问题。

Rapid progress in machine learning and artificial intelligence (AI) has brought increasing attention to the potential impacts of AI technologies on society. In this paper we discuss one such potential impact: the problem of accidents in machine learning systems, defined as unintended and harmful behavior that may emerge from poor design of real-world AI systems. We present a list of five practical research problems related to accident risk, categorized according to whether the problem originates from having the wrong objective function ("avoiding side effects" and "avoiding reward hacking"), an objective function that is too expensive to evaluate frequently ("scalable supervision"), or undesirable behavior during the learning process ("safe exploration" and "distributional shift"). We review previous work in these areas as well as suggesting research directions with a focus on relevance to cutting-edge AI systems. Finally, we consider the high-level question of how to think most productively about the safety of forward-looking applications of AI.

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 11)

阅读逐段中英对照全文 →