通过增强弱专家来监督强学习者

Supervising strong learners by amplifying weak experts

达里奥·阿莫迪 Dario Amodei · OpenAI · 2018-10-19 · arXiv:1810.08575 ↗ · 被引 174

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

许多现实世界的学习任务涉及复杂或难以指定的目标,使用更容易指定的代理目标可能导致性能不佳或行为失调。一种解决方案是由人类通过演示或判断表现来提供训练信号,但如果任务过于复杂以至于人类无法直接评估,这种方法就会失败。我们提出了迭代放大,这是一种替代训练策略,通过组合较易子问题的解决方案,逐步为困难问题构建训练信号。迭代放大与专家迭代(Anthony 等人,2017;Silver 等人,2017)密切相关,区别在于它不使用外部奖励函数。我们在算法环境中展示了结果,表明迭代放大可以高效地学习复杂行为。

Many real world learning tasks involve complex or hard-to-specify objectives, and using an easier-to-specify proxy can lead to poor performance or misaligned behavior. One solution is to have humans provide a training signal by demonstrating or judging performance, but this approach fails if the task is too complicated for a human to directly evaluate. We propose Iterated Amplification, an alternative training strategy which progressively builds up a training signal for difficult problems by combining solutions to easier subproblems. Iterated Amplification is closely related to Expert Iteration (Anthony et al., 2017; Silver et al., 2017), except that it uses no external reward function. We present results in algorithmic environments, showing that Iterated Amplification can efficiently learn complex behaviors.

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 15)

阅读逐段中英对照全文 →