AlphaGeometry: An Olympiad-level AI system for geometry
打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→AlphaGeometry 是一个神经符号系统,由神经语言模型和符号推理引擎组成,两者协同工作,为复杂的几何定理寻找证明。类似于“思考,快与慢”的理念,一个系统提供快速、“直觉”的想法,而另一个则提供更审慎、理性的决策。 由于语言模型擅长识别数据中的一般模式和关系,它们能快速预测可能有用的构造,但往往缺乏严格推理或解释其决策的能力。另一方面,符号推理引擎基于形式逻辑,使用明确的规则得出结论。它们是理性且可解释的,但可能“缓慢”且不灵活——尤其是在独自处理大型复杂问题时。 AlphaGeometry 的语言模型引导其符号推理引擎朝着几何问题的可能解决方案前进。奥林匹克几何问题基于图形,需要添加新的几何构造(如点、线或圆)才能求解。AlphaGeometry 的语言模型从无限的可能性中预测哪些新构造最有用。这些线索有助于填补空白,使符号引擎能够对图形进行进一步推理,并逐步接近解决方案。
AlphaGeometry is a neuro-symbolic system made up of a neural language model and a symbolic deduction engine, which work together to find proofs for complex geometry theorems. Akin to the idea of “thinking, fast and slow”, one system provides fast, “intuitive” ideas, and the other, more deliberate, rational decision-making. Because language models excel at identifying general patterns and relationships in data, they can quickly predict potentially useful constructs, but often lack the ability to reason rigorously or explain their decisions. Symbolic deduction engines, on the other hand, are based on formal logic and use clear rules to arrive at conclusions. They are rational and explainable, but they can be “slow” and inflexible - especially when dealing with large, complex problems on their own.
AlphaGeometry 是一个神经符号系统,由神经语言模型和符号推理引擎组成,两者协同工作,为复杂的几何定理寻找证明。类似于“思考,快与慢”的理念,一个系统提供快速、“直觉”的想法,而另一个则提供更审慎、理性的决策。
AlphaGeometry is a neuro-symbolic system made up of a neural language model and a symbolic deduction engine, which work together to find proofs for complex geometry theorems. Akin to the idea of “thinking, fast and slow”, one system provides fast, “intuitive” ideas, and the other, more deliberate, rational decision-making.
由于语言模型擅长识别数据中的一般模式和关系,它们能快速预测可能有用的构造,但往往缺乏严格推理或解释其决策的能力。另一方面,符号推理引擎基于形式逻辑,使用明确的规则得出结论。它们是理性且可解释的,但可能“缓慢”且不灵活——尤其是在独自处理大型复杂问题时。
Because language models excel at identifying general patterns and relationships in data, they can quickly predict potentially useful constructs, but often lack the ability to reason rigorously or explain their decisions. Symbolic deduction engines, on the other hand, are based on formal logic and use clear rules to arrive at conclusions. They are rational and explainable, but they can be “slow” and inflexible - especially when dealing with large, complex problems on their own.
AlphaGeometry 的语言模型引导其符号推理引擎朝着几何问题的可能解决方案前进。奥林匹克几何问题基于图形,需要添加新的几何构造(如点、线或圆)才能求解。AlphaGeometry 的语言模型从无限的可能性中预测哪些新构造最有用。这些线索有助于填补空白,使符号引擎能够对图形进行进一步推理,并逐步接近解决方案。
AlphaGeometry’s language model guides its symbolic deduction engine towards likely solutions to geometry problems. Olympiad geometry problems are based on diagrams that need new geometric constructs to be added before they can be solved, such as points, lines or circles. AlphaGeometry’s language model predicts which new constructs would be most useful to add, from an infinite number of possibilities. These clues help fill in the gaps and allow the symbolic engine to make further deductions about the diagram and close in on the solution.
AlphaGeometry 解决一个简单问题:给定问题图形及其定理前提(左),AlphaGeometry(中)首先使用其符号引擎推导关于图形的新陈述,直到找到解决方案或穷尽新陈述。如果未找到解决方案,AlphaGeometry 的语言模型会添加一个可能有用的构造(蓝色),为符号引擎开辟新的推理路径。此循环持续进行,直到找到解决方案(右)。在此示例中,仅需一个构造。
AlphaGeometry solving a simple problem: Given the problem diagram and its theorem premises (left), AlphaGeometry (middle) first uses its symbolic engine to deduce new statements about the diagram until the solution is found or new statements are exhausted. If no solution is found, AlphaGeometry’s language model adds one potentially useful construct (blue), opening new paths of deduction for the symbolic engine. This loop continues until a solution is found (right). In this example, just one construct is required.
AlphaGeometry 解决一个奥林匹克问题:2015 年国际数学奥林匹克竞赛第 3 题(左)和 AlphaGeometry 解决方案的压缩版本(右)。蓝色元素是添加的构造。AlphaGeometry 的解决方案包含 109 个逻辑步骤。
AlphaGeometry solving an Olympiad problem: Problem 3 of the 2015 International Mathematics Olympiad (left) and a condensed version of AlphaGeometry’s solution (right). The blue elements are added constructs. AlphaGeometry’s solution has 109 logical steps.
几何学依赖于对空间、距离、形状和相对位置的理解,是艺术、建筑、工程及许多其他领域的基础。人类可以用纸笔学习几何,通过检查图形并利用现有知识来发现新的、更复杂的几何性质和关系。我们的合成数据生成方法大规模模拟了这一知识构建过程,使我们能够从头训练 AlphaGeometry,无需任何人类演示。
Geometry relies on understanding of space, distance, shape, and relative positions, and is fundamental to art, architecture, engineering and many other fields. Humans can learn geometry using a pen and paper, examining diagrams and using existing knowledge to uncover new, more sophisticated geometric properties and relationships. Our synthetic data generation approach emulates this knowledge-building process at scale, allowing us to train AlphaGeometry from scratch, without any human demonstrations.
利用高度并行化的计算,系统首先生成了十亿个随机几何图形,并详尽推导出每个图形中所有点与线之间的关系。AlphaGeometry 找到了每个图形中包含的所有证明,然后反向推导出需要哪些额外的构造(如果有的话)才能得到这些证明。我们将这一过程称为“符号演绎与回溯”。
Using highly parallelized computing, the system started by generating one billion random diagrams of geometric objects and exhaustively derived all the relationships between the points and lines in each diagram. AlphaGeometry found all the proofs contained in each diagram, then worked backwards to find out what additional constructs, if any, were needed to arrive at those proofs. We call this process “symbolic deduction and traceback”.
AlphaGeometry 生成的合成数据的可视化表示
Visual representations of the synthetic data generated by AlphaGeometry
这个庞大的数据池经过过滤,排除了相似的示例,最终得到了一个包含 1 亿个不同难度的独特示例的训练数据集,其中 900 万个示例包含添加的构造。凭借如此多的关于这些构造如何导致证明的示例,AlphaGeometry 的语言模型能够在遇到奥林匹克几何问题时,为新的构造提出良好的建议。
That huge data pool was filtered to exclude similar examples, resulting in a final training dataset of 100 million unique examples of varying difficulty, of which nine million featured added constructs. With so many examples of how these constructs led to proofs, AlphaGeometry’s language model is able to make good suggestions for new constructs when presented with Olympiad geometry problems.
AlphaGeometry 提供的每个奥数题解都经过计算机检查和验证。我们还将它的结果与之前的 AI 方法以及人类在奥赛中的表现进行了比较。此外,数学教练、前奥赛金牌得主 Evan Chen 为我们评估了 AlphaGeometry 的部分解答。
The solution to every Olympiad problem provided by AlphaGeometry was checked and verified by computer. We also compared its results with previous AI methods, and with human performance at the Olympiad. In addition, Evan Chen, a math coach and former Olympiad gold-medalist, evaluated a selection of AlphaGeometry’s solutions for us.
Chen 说:“AlphaGeometry 的输出令人印象深刻,因为它既可验证又简洁。过去基于证明的竞赛题的 AI 解决方案有时好坏参半(输出有时正确,需要人工检查)。AlphaGeometry 没有这个弱点:它的解具有机器可验证的结构。尽管如此,它的输出仍然可读。人们可能想象过一种通过暴力坐标系统解决几何问题的计算机程序:想想一页又一页繁琐的代数计算。AlphaGeometry 不是这样。它使用经典的几何规则,涉及角度和相似三角形,就像学生做的那样。”
Chen said: “AlphaGeometry's output is impressive because it's both verifiable and clean. Past AI solutions to proof-based competition problems have sometimes been hit-or-miss (outputs are only correct sometimes and need human checks). AlphaGeometry doesn't have this weakness: its solutions have machine-verifiable structure. Yet despite this, its output is still human-readable. One could have imagined a computer program that solved geometry problems by brute-force coordinate systems: think pages and pages of tedious algebra calculation. AlphaGeometry is not that. It uses classical geometry rules with angles and similar triangles just as students do.”
由于每届奥赛有六道题,通常只有两道是几何题,因此 AlphaGeometry 只能应用于给定奥赛的三分之一题目。尽管如此,仅凭其几何能力,它就成为了世界上第一个能够在 2000 年和 2015 年达到 IMO 铜牌门槛的 AI 模型。
As each Olympiad features six problems, only two of which are typically focused on geometry, AlphaGeometry can only be applied to one-third of the problems at a given Olympiad. Nevertheless, its geometry capability alone makes it the first AI model in the world capable of passing the bronze medal threshold of the IMO in 2000 and 2015.
在几何方面,我们的系统接近 IMO 金牌得主的水平,但我们着眼于更大的目标:推进下一代 AI 系统的推理能力。考虑到使用大规模合成数据从头训练 AI 系统的更广泛潜力,这种方法可能会塑造未来的 AI 系统在数学及其他领域发现新知识的方式。
In geometry, our system approaches the standard of an IMO gold-medalist, but we have our eye on an even bigger prize: advancing reasoning for next-generation AI systems. Given the wider potential of training AI systems from scratch with large-scale synthetic data, this approach could shape how the AI systems of the future discover new knowledge, in math and beyond.
AlphaGeometry 建立在 Google DeepMind 和 Google Research 用 AI 开创数学推理先河的工作之上——从探索纯数学之美到用语言模型解决数学和科学问题。最近,我们推出了 FunSearch,它使用大型语言模型在数学科学的开放问题上取得了首次发现。
AlphaGeometry builds on Google DeepMind and Google Research’s work to pioneer mathematical reasoning with AI – from exploring the beauty of pure mathematics to solving mathematical and scientific problems with language models. And most recently, we introduced FunSearch, which made the first discoveries in open problems in mathematical sciences using Large Language Models.
我们的长期目标仍然是构建能够在数学领域泛化的 AI 系统,开发通用 AI 系统所依赖的复杂问题解决和推理能力,同时拓展人类知识的边界。
Our long-term goal remains to build AI systems that can generalize across mathematical fields, developing the sophisticated problem-solving and reasoning that general AI systems will depend on, all the while extending the frontiers of human knowledge.
阅读我们在《自然》杂志上的论文访问 AlphaGeometry GitHub
Read our paper in NatureAccess the AlphaGeometry Github
本项目是 Google DeepMind 团队与纽约大学计算机科学系的合作成果。本工作的作者包括 Trieu Trinh、Yuhuai Wu、Quoc Le、He He 和 Thang Luong。我们感谢 Rif A. Saurous、Denny Zhou、Christian Szegedy、Delesley Hutchins、Thomas Kipf、Hieu Pham、Petar Veličković、Edward Lockhart、Debidatta Dwibedi、Kyunghyun Cho、Lerrel Pinto、Alfredo Canziani、Thomas Wies、He He 的研究团队、Evan Chen、Mirek Olsak、Patrik Bak 的帮助和支持。我们还要感谢 Google DeepMind 领导层的支持,特别是 Ed Chi、Koray Kavukcuoglu、Pushmeet Kohli 和 Demis Hassabis。
This project is a collaboration between the Google DeepMind team and the Computer Science Department of New York University. The authors of this work include Trieu Trinh, Yuhuai Wu, Quoc Le, He He, and Thang Luong. We thank Rif A. Saurous, Denny Zhou, Christian Szegedy, Delesley Hutchins, Thomas Kipf, Hieu Pham, Petar Veličković, Edward Lockhart, Debidatta Dwibedi, Kyunghyun Cho, Lerrel Pinto, Alfredo Canziani, Thomas Wies, He He’s research group, Evan Chen, Mirek Olsak, Patrik Bak for their help and support. We would also like to thank Google DeepMind leadership for the support, especially Ed Chi, Koray Kavukcuoglu, Pushmeet Kohli, and Demis Hassabis.