强化学习的分布视角

A Distributional Perspective on Reinforcement Learning

马克·贝尔马尔 Marc G. Bellemare · Google DeepMind · 2017-07-21 · arXiv:1707.06887 ↗ · 被引 1882

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

本文论证了价值分布(强化学习智能体获得的随机回报的分布)的根本重要性,这与通常只建模回报期望(即价值)的强化学习方法形成对比。尽管已有文献研究价值分布,但此前它总是用于特定目的,如实现风险感知行为。我们从策略评估和控制设置的理论结果开始,揭示了后者中显著的分布不稳定性。然后利用分布视角设计了一种新算法,将贝尔曼方程应用于近似价值分布的学习。我们使用街机学习环境中的游戏套件评估了该算法,获得了最先进的结果和轶事证据,证明了价值分布在近似强化学习中的重要性。最后,我们结合理论和实证证据,强调了价值分布在近似设置中影响学习的方式。

In this paper we argue for the fundamental importance of the value distribution: the distribution of the random return received by a reinforcement learning agent. This is in contrast to the common approach to reinforcement learning which models the expectation of this return, or value. Although there is an established body of literature studying the value distribution, thus far it has always been used for a specific purpose such as implementing risk-aware behaviour. We begin with theoretical results in both the policy evaluation and control settings, exposing a significant distributional instability in the latter. We then use the distributional perspective to design a new algorithm which applies Bellman's equation to the learning of approximate value distributions. We evaluate our algorithm using the suite of games from the Arcade Learning Environment. We obtain both state-of-the-art results and anecdotal evidence demonstrating the importance of the value distribution in approximate reinforcement learning. Finally, we combine theoretical and empirical evidence to highlight the ways in which the value distribution impacts learning in the approximate setting.

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 19)

阅读逐段中英对照全文 →