通过通用强化学习算法的自我对弈掌握国际象棋和将棋

Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

扬尼斯·安东诺格鲁 Ioannis Antonoglou · Google DeepMind · 2017-12-05 · arXiv:1712.01815 ↗ · 被引 2101

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

国际象棋是人工智能历史上研究最广泛的领域。最强的程序基于复杂搜索技术、领域特定调整和人类专家数十年来精心设计的评估函数的结合。相比之下,AlphaGo Zero 程序最近通过从自我对弈游戏中进行的无先验知识强化学习,在围棋中实现了超人类表现。在本文中,我们将这种方法推广为一个单一的 AlphaZero 算法,该算法可以在许多具有挑战性的领域中通过无先验知识学习达到超人类表现。从随机对弈开始,除了游戏规则外没有任何领域知识,AlphaZero 在 24 小时内就在国际象棋、将棋(日本象棋)以及围棋中达到了超人类水平,并在每种情况下令人信服地击败了世界冠军程序。

The game of chess is the most widely-studied domain in the history of artificial intelligence. The strongest programs are based on a combination of sophisticated search techniques, domain-specific adaptations, and handcrafted evaluation functions that have been refined by human experts over several decades. In contrast, the AlphaGo Zero program recently achieved superhuman performance in the game of Go, by tabula rasa reinforcement learning from games of self-play. In this paper, we generalise this approach into a single AlphaZero algorithm that can achieve, tabula rasa, superhuman performance in many challenging domains. Starting from random play, and given no domain knowledge except the game rules, AlphaZero achieved within 24 hours a superhuman level of play in the games of chess and shogi (Japanese chess) as well as Go, and convincingly defeated a world-champion program in each case.

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 1)

阅读逐段中英对照全文 →