世界模型

World Models

戴维·哈 David Ha · Google · 2018-03-27 · arXiv:1803.10122 ↗ · 被引 1858

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

我们探索构建流行强化学习环境的生成式神经网络模型。我们的世界模型可以通过无监督方式快速训练,学习环境的压缩时空表征。通过将世界模型提取的特征作为智能体的输入,我们可以训练一个非常紧凑且简单的策略来解决所需任务。我们甚至可以在世界模型生成的幻觉梦境中完全训练智能体,并将该策略迁移回实际环境。本文的交互式版本可在 https://worldmodels.github.io/ 获取。

We explore building generative neural network models of popular reinforcement learning environments. Our world model can be trained quickly in an unsupervised manner to learn a compressed spatial and temporal representation of the environment. By using features extracted from the world model as inputs to an agent, we can train a very compact and simple policy that can solve the required task. We can even train our agent entirely inside of its own hallucinated dream generated by its world model, and transfer this policy back into the actual environment. An interactive version of this paper is available at https://worldmodels.github.io/

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 27)

阅读逐段中英对照全文 →