Through multi-agent competition, the simple objective of hide-and-seek, and standard reinforcement learning algorithms at scale, we find that agents create a self-supervised autocurriculum inducing multiple distinct rounds of emergent strategy, many of which require sophisticated tool use and coordination. We find clear evidence of six emergent phases in agent strategy in our environment, each of which creates a new pressure for the opposing team to adapt; for instance, agents learn to build multi-object shelters using moveable boxes which in turn leads to agents discovering that they can overcome obstacles using ramps. We further provide evidence that multi-agent competition may scale better with increasing environment complexity and leads to behavior that centers around far more human-relevant skills than other self-supervised reinforcement learning methods such as intrinsic motivation. Finally, we propose transfer and fine-tuning as a way to quantitatively evaluate targeted capabilities, and we compare hide-and-seek agents to both intrinsic motivation and random initialization baselines in a suite of domain-specific intelligence tests.
核心贡献 · Key contributions
通过多智能体自动课程展示了捉迷藏中的六个涌现策略阶段。 Demonstrates six emergent phases of strategy in hide-and-seek via multi-agent autocurricula.
表明多智能体竞争导致复杂的工具使用和协调。 Shows that multi-agent competition leads to sophisticated tool use and coordination.
提供证据表明多智能体自动课程在环境复杂性上比内在动机扩展得更好。 Provides evidence that multi-agent autocurricula scale better with environment complexity than intrinsic motivation.
提出迁移学习和微调作为目标能力的定量评估方法。 Proposes transfer learning and fine-tuning as a quantitative evaluation method for targeted capabilities.
引入一套领域特定的智能测试来评估智能体能力。 Introduces a suite of domain-specific intelligence tests for evaluating agent capabilities.
开源环境和代码以鼓励在物理基础的多智能体自动课程方面的进一步研究。 Open-sources environments and code to encourage further research in physically grounded multi-agent autocurricula.
局限 · Limitations
策略空间固有有限,可能不会超过六种涌现模式。 Strategy space is inherently bounded and may not surpass six emergent modes.
由于奖励函数不对齐,智能体需要大量经验才能通过阶段。 Agents require enormous experience to progress through stages due to misaligned reward functions.
由于技能表示纠缠且难以微调,迁移结果参差不齐。 Transfer results are mixed due to entangled skill representations that are hard to fine-tune.
智能体利用环境设计的不准确性,例如物理模拟缺陷。 Agents exploit environment design inaccuracies, such as physics simulation flaws.
样本复杂度高;降低它是未来研究的重要方向。 Sample complexity is high; reducing it is an important direction for future research.
论文章节 · Sections(共 8)
多智能体自动课程中的涌现工具使用Emergent Tool Use From Multi-Agent Autocurricula