通过世界模型掌握多样领域

Mastering Diverse Domains through World Models

达尼亚尔·哈夫纳 Danijar Hafner · Google DeepMind / University of Toronto · 2023-01-10 · arXiv:2301.04104 ↗ · 被引 1228

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

开发一种能够学习解决广泛领域任务的通用算法一直是人工智能的基本挑战。尽管当前的强化学习算法可以轻松应用于与其开发环境相似的任务,但将它们配置到新的应用领域需要大量的人类专业知识和实验。我们提出了 DreamerV3,这是一种通用算法,通过单一配置在超过 150 个不同任务上超越了专门方法。Dreamer 学习环境模型,并通过想象未来场景来改进其行为。基于归一化、平衡和变换的鲁棒性技术使其能够在不同领域稳定学习。开箱即用,Dreamer 是第一个从零开始在 Minecraft 中收集钻石的算法,无需人类数据或课程。这一成就被认为是人工智能中的重大挑战,需要从像素和稀疏奖励中探索远见策略。我们的工作使得无需大量实验即可解决具有挑战性的控制问题,从而使得强化学习具有广泛的适用性。

Developing a general algorithm that learns to solve tasks across a wide range of applications has been a fundamental challenge in artificial intelligence. Although current reinforcement learning algorithms can be readily applied to tasks similar to what they have been developed for, configuring them for new application domains requires significant human expertise and experimentation. We present DreamerV3, a general algorithm that outperforms specialized methods across over 150 diverse tasks, with a single configuration. Dreamer learns a model of the environment and improves its behavior by imagining future scenarios. Robustness techniques based on normalization, balancing, and transformations enable stable learning across domains. Applied out of the box, Dreamer is the first algorithm to collect diamonds in Minecraft from scratch without human data or curricula. This achievement has been posed as a significant challenge in artificial intelligence that requires exploring farsighted strategies from pixels and sparse rewards in an open world. Our work allows solving challenging control problems without extensive experimentation, making reinforcement learning broadly applicable.

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 9)

阅读逐段中英对照全文 →