多智能体自动课程中的涌现工具使用

Emergent Tool Use From Multi-Agent Autocurricula

OpenAI OpenAI · OpenAI · 2019-09-17 · arXiv:1909.07528 ↗ · 被引 765

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

通过多智能体竞争、简单的捉迷藏目标以及标准强化学习算法的大规模应用,我们发现智能体创造了一个自监督的自动课程,诱导出多轮不同的涌现策略,其中许多需要复杂的工具使用和协调。我们在环境中发现了智能体策略的六个涌现阶段的明确证据,每个阶段都为对手团队创造了新的适应压力;例如,智能体学会使用可移动箱子建造多物体掩体,这进而导致智能体发现可以利用斜坡克服障碍。我们进一步提供证据表明,多智能体竞争可能随着环境复杂性的增加而更好地扩展,并导致行为集中在比其他自监督强化学习方法(如内在动机)更与人类相关的技能上。最后,我们提出迁移和微调作为定量评估目标能力的方法,并在领域特定的智能测试套件中将捉迷藏智能体与内在动机和随机初始化基线进行比较。

Through multi-agent competition, the simple objective of hide-and-seek, and standard reinforcement learning algorithms at scale, we find that agents create a self-supervised autocurriculum inducing multiple distinct rounds of emergent strategy, many of which require sophisticated tool use and coordination. We find clear evidence of six emergent phases in agent strategy in our environment, each of which creates a new pressure for the opposing team to adapt; for instance, agents learn to build multi-object shelters using moveable boxes which in turn leads to agents discovering that they can overcome obstacles using ramps. We further provide evidence that multi-agent competition may scale better with increasing environment complexity and leads to behavior that centers around far more human-relevant skills than other self-supervised reinforcement learning methods such as intrinsic motivation. Finally, we propose transfer and fine-tuning as a way to quantitatively evaluate targeted capabilities, and we compare hide-and-seek agents to both intrinsic motivation and random initialization baselines in a suite of domain-specific intelligence tests.

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 8)

阅读逐段中英对照全文 →