梦想控制:通过潜在想象学习行为

Dream to Control: Learning Behaviors by Latent Imagination

达尼亚尔·哈夫纳 Danijar Hafner · Google DeepMind · 2019-12-03 · arXiv:1912.01603 ↗ · 被引 2052

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

学到的世界模型总结了智能体的经验,以促进复杂行为的学习。虽然通过深度学习从高维感官输入中学习世界模型正变得可行,但从中推导行为有许多潜在方法。我们提出了 Dreamer,一个仅通过潜在想象从图像中解决长时域任务的强化学习智能体。我们通过将学到的状态值的解析梯度反向传播通过在一个学到的世界模型的紧凑状态空间中想象的轨迹,来高效地学习行为。在 20 个具有挑战性的视觉控制任务上,Dreamer 在数据效率、计算时间和最终性能方面超越了现有方法。

Learned world models summarize an agent's experience to facilitate learning complex behaviors. While learning world models from high-dimensional sensory inputs is becoming feasible through deep learning, there are many potential ways for deriving behaviors from them. We present Dreamer, a reinforcement learning agent that solves long-horizon tasks from images purely by latent imagination. We efficiently learn behaviors by propagating analytic gradients of learned state values back through trajectories imagined in the compact state space of a learned world model. On 20 challenging visual control tasks, Dreamer exceeds existing approaches in data-efficiency, computation time, and final performance.

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 29)

阅读逐段中英对照全文 →