彩虹:融合深度强化学习的多项改进

Rainbow: Combining Improvements in Deep Reinforcement Learning

哈多·范哈塞尔特 Hado van Hasselt · Google DeepMind · 2017-10-06 · arXiv:1710.02298 ↗ · 被引 2642

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

深度强化学习社区对 DQN 算法进行了多项独立的改进。然而,目前尚不清楚这些扩展中哪些是互补的,可以有效地结合起来。本文研究了 DQN 算法的六种扩展,并实证研究了它们的组合效果。我们的实验表明,这种组合在 Atari 2600 基准测试中,无论是在数据效率还是最终性能方面,都达到了最先进的水平。我们还提供了详细的消融研究结果,展示了每个组成部分对整体性能的贡献。

The deep reinforcement learning community has made several independent improvements to the DQN algorithm. However, it is unclear which of these extensions are complementary and can be fruitfully combined. This paper examines six extensions to the DQN algorithm and empirically studies their combination. Our experiments show that the combination provides state-of-the-art performance on the Atari 2600 benchmark, both in terms of data efficiency and final performance. We also provide results from a detailed ablation study that shows the contribution of each component to overall performance.

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 21)

阅读逐段中英对照全文 →