AI achieves silver-medal standard solving International Mathematical Olympiad problems
打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→IMO 是历史最悠久、规模最大、最具声望的青年数学家竞赛,自 1959 年以来每年举办。 每年,顶尖的大学预科数学学生要训练数千小时,以解决代数、组合学、几何和数论中六道极其困难的问题。许多菲尔兹奖(数学最高荣誉之一)得主都曾代表其国家参加 IMO。 近年来,年度 IMO 竞赛也被广泛认为是机器学习领域的一项重大挑战,以及衡量 AI 系统高级数学推理能力的理想基准。
The IMO is the oldest, largest and most prestigious competition for young mathematicians, held annually since 1959. Each year, elite pre-college mathematicians train, sometimes for thousands of hours, to solve six exceptionally difficult problems in algebra, combinatorics, geometry and number theory. Many of the winners of the Fields Medal, one of the highest honors for mathematicians, have represented their country at the IMO. More recently, the annual IMO competition has also become widely recognised as a grand challenge in machine learning and an aspirational benchmark for measuring an AI system’s advanced mathematical reasoning capabilities.
IMO 是历史最悠久、规模最大、最具声望的青年数学家竞赛,自 1959 年以来每年举办。
The IMO is the oldest, largest and most prestigious competition for young mathematicians, held annually since 1959.
每年,顶尖的大学预科数学学生要训练数千小时,以解决代数、组合学、几何和数论中六道极其困难的问题。许多菲尔兹奖(数学最高荣誉之一)得主都曾代表其国家参加 IMO。
Each year, elite pre-college mathematicians train, sometimes for thousands of hours, to solve six exceptionally difficult problems in algebra, combinatorics, geometry and number theory. Many of the winners of the Fields Medal, one of the highest honors for mathematicians, have represented their country at the IMO.
近年来,年度 IMO 竞赛也被广泛认为是机器学习领域的一项重大挑战,以及衡量 AI 系统高级数学推理能力的理想基准。
More recently, the annual IMO competition has also become widely recognised as a grand challenge in machine learning and an aspirational benchmark for measuring an AI system’s advanced mathematical reasoning capabilities.
今年,我们将组合 AI 系统应用于 IMO 组织者提供的竞赛问题。我们的解决方案由著名数学家、IMO 金牌得主和菲尔兹奖得主 Timothy Gowers 爵士教授以及两次 IMO 金牌得主、IMO 2024 问题遴选委员会主席 Joseph Myers 博士根据 IMO 的评分规则进行评分。
This year, we applied our combined AI system to the competition problems, provided by the IMO organizers. Our solutions were scored according to the IMO’s point-awarding rules by prominent mathematicians Prof Sir Timothy Gowers, an IMO gold medalist and Fields Medal winner, and Dr Joseph Myers, a two-time IMO gold medalist and Chair of the IMO 2024 Problem Selection Committee.
IMO 金牌得主和菲尔兹奖得主
IMO gold medalist and Fields Medal winner
首先,问题被手动翻译成形式化数学语言,以便我们的系统理解。在正式比赛中,学生分两次各 4.5 小时提交答案。我们的系统在几分钟内解决了一个问题,而解决其他问题则用了最多三天。
First, the problems were manually translated into formal mathematical language for our systems to understand. In the official competition, students submit answers in two sessions of 4.5 hours each. Our systems solved one problem within minutes and took up to three days to solve the others.
AlphaProof 通过确定答案并证明其正确性,解决了两个代数问题和一个数论问题。这包括竞赛中最难的问题,今年 IMO 只有五名参赛者解出。AlphaGeometry 2 证明了几何问题,而两个组合问题仍未解决。
AlphaProof solved two algebra problems and one number theory problem by determining the answer and proving it was correct. This included the hardest problem in the competition, solved by only five contestants at this year’s IMO. AlphaGeometry 2 proved the geometry problem, while the two combinatorics problems remained unsolved.
每个问题最高可得 7 分,总分 42 分。我们的系统最终获得 28 分,在每个已解决问题上均获得满分——相当于银牌类别的高分。今年,金牌门槛为 29 分,在官方竞赛的 609 名参赛者中有 58 人达到。
Each of the six problems can earn seven points, with a total maximum of 42. Our system achieved a final score of 28 points, earning a perfect score on each problem solved — equivalent to the top end of the silver-medal category. This year, the gold-medal threshold starts at 29 points, and was achieved by 58 of 609 contestants at the official competition.
图表展示了我们的 AI 系统与 IMO 2024 人类参赛者相比的表现。我们获得 42 分中的 28 分,达到了与银牌得主相同的水平。
Graph showing performance of our AI system relative to human competitors at IMO 2024. We earned 28 out of 42 total points, achieving the same level as a silver medalist in the competition.
AlphaProof 是一个在形式语言 Lean 中自我训练以证明数学陈述的系统。它结合了一个预训练语言模型与 AlphaZero 强化学习算法,后者此前已自学掌握了国际象棋、将棋和围棋。
AlphaProof is a system that trains itself to prove mathematical statements in the formal language Lean. It couples a pre-trained language model with the AlphaZero reinforcement learning algorithm, which previously taught itself how to master the games of chess, shogi and Go.
形式语言的关键优势在于,涉及数学推理的证明可以被形式化地验证其正确性。然而,它们在机器学习中的使用此前受到可用人类编写数据量极少的限制。
Formal languages offer the critical advantage that proofs involving mathematical reasoning can be formally verified for correctness. Their use in machine learning has, however, previously been constrained by the very limited amount of human-written data available.
相比之下,基于自然语言的方法尽管可以访问数量级更多的数据,却可能产生看似合理但实际上不正确的中间推理步骤和解决方案。我们通过微调 Gemini 模型来自动将自然语言问题陈述翻译成形式化陈述,从而在这两个互补领域之间建立了一座桥梁,创建了一个包含不同难度形式化问题的大型库。
In contrast, natural language based approaches can hallucinate plausible but incorrect intermediate reasoning steps and solutions, despite having access to orders of magnitudes more data. We established a bridge between these two complementary spheres by fine-tuning a Gemini model to automatically translate natural language problem statements into formal statements, creating a large library of formal problems of varying difficulty.
当面对一个问题时,AlphaProof 会生成候选解决方案,然后通过在 Lean 中搜索可能的证明步骤来证明或反驳它们。每个被发现并验证的证明都用于强化 AlphaProof 的语言模型,增强其解决后续更具挑战性问题的能力。
When presented with a problem, AlphaProof generates solution candidates and then proves or disproves them by searching over possible proof steps in Lean. Each proof that was found and verified is used to reinforce AlphaProof’s language model, enhancing its ability to solve subsequent, more challenging problems.
我们通过证明或反驳数百万个问题来训练 AlphaProof 参加国际数学奥林匹克竞赛,这些问题涵盖了广泛的难度和数学主题领域,训练持续了比赛前的数周时间。训练循环也在比赛期间应用,通过强化对比赛问题自生成变体的证明,直到找到完整的解决方案。
We trained AlphaProof for the IMO by proving or disproving millions of problems, covering a wide range of difficulties and mathematical topic areas over a period of weeks leading up to the competition. The training loop was also applied during the contest, reinforcing proofs of self-generated variations of the contest problems until a full solution could be found.
AlphaProof 强化学习训练循环的过程信息图:大约一百万个非正式数学问题通过形式化器网络被翻译成形式化数学语言。然后,求解器网络搜索这些问题的证明或反例,并通过 AlphaZero 算法逐步自我训练以解决更具挑战性的问题。
Process infographic of AlphaProof’s reinforcement learning training loop: Around one million informal math problems are translated into a formal math language by a formalizer network. Then a solver network searches for proofs or disproofs of the problems, progressively training itself via the AlphaZero algorithm to solve more challenging problems.
AlphaGeometry 2 是 AlphaGeometry 的显著改进版本。它是一个神经符号混合系统,其中语言模型基于 Gemini,并在比其前身多一个数量级的合成数据上从头训练。这帮助模型处理更具挑战性的几何问题,包括涉及物体运动以及角度、比例或距离方程的问题。
AlphaGeometry 2 is a significantly improved version of AlphaGeometry. It’s a neuro-symbolic hybrid system in which the language model was based on Gemini and trained from scratch on an order of magnitude more synthetic data than its predecessor. This helped the model tackle much more challenging geometry problems, including problems about movements of objects and equations of angles, ratio or distances.
AlphaGeometry 2 采用了一个比其前身快两个数量级的符号引擎。当遇到新问题时,使用一种新颖的知识共享机制,使得不同搜索树的高级组合能够解决更复杂的问题。
AlphaGeometry 2 employs a symbolic engine that is two orders of magnitude faster than its predecessor. When presented with a new problem, a novel knowledge-sharing mechanism is used to enable advanced combinations of different search trees to tackle more complex problems.
在今年竞赛之前,AlphaGeometry 2 能够解决过去 25 年所有历史 IMO 几何问题的 83%,而前身的解决率为 53%。对于 IMO 2024,AlphaGeometry 2 在收到形式化后 19 秒内解决了问题 4。
Before this year’s competition, AlphaGeometry 2 could solve 83% of all historical IMO geometry problems from the past 25 years, compared to the 53% rate achieved by its predecessor. For IMO 2024, AlphaGeometry 2 solved Problem 4 within 19 seconds after receiving its formalization.
问题 4 的图示,要求证明 ∠KIL 和 ∠XPY 之和为 180°。AlphaGeometry 2 提出构造点 E,位于直线 BI 上使得 ∠AEB = 90°。点 E 有助于赋予 AB 中点 L 以意义,创建了许多相似三角形对,如 ABE ∼ YBI 和 ALE ∼ IPC,这些是证明结论所需的。
Illustration of Problem 4, which asks to prove the sum of ∠KIL and ∠XPY equals 180°. AlphaGeometry 2 proposed to construct E, a point on the line BI so that ∠AEB = 90°. Point E helps give purpose to the midpoint L of AB, creating many pairs of similar triangles such as ABE ~ YBI and ALE ~ IPC needed to prove the conclusion.
作为我们国际数学奥林匹克(IMO)工作的一部分,我们还试验了一个自然语言推理系统,该系统基于 Gemini 和我们最新的研究,以实现高级问题解决能力。该系统不需要将问题翻译成形式语言,并且可以与其他 AI 系统结合。我们还在今年的 IMO 问题上测试了这种方法,结果显示前景广阔。
As part of our IMO work, we also experimented with a natural language reasoning system, built upon Gemini and our latest research to enable advanced problem-solving skills. This system doesn’t require the problems to be translated into a formal language and could be combined with other AI systems. We also tested this approach on this year’s IMO problems and the results showed great promise.
我们的团队正在继续探索多种 AI 方法来推进数学推理,我们已在《自然》杂志的一篇论文中发布了关于 AlphaProof 的更多技术细节。
Our teams are continuing to explore multiple AI approaches for advancing mathematical reasoning and we have released more technical details on AlphaProof in a Nature paper.
我们对这样一个未来感到兴奋:数学家与 AI 工具合作,探索假设,尝试解决长期问题的全新方法,并快速完成证明中耗时的元素——同时像 Gemini 这样的 AI 系统在数学和更广泛的推理方面变得更加有能力。
We’re excited for a future in which mathematicians work with AI tools to explore hypotheses, try bold new approaches to solving long-standing problems and quickly complete time-consuming elements of proofs — and where AI systems like Gemini become more capable at math and broader reasoning.
我们感谢国际数学奥林匹克组织的支持。
We thank the International Mathematical Olympiad organization for their support.
AlphaProof 的开发由 Thomas Hubert、Rishi Mehta 和 Laurent Sartran 领导;AlphaGeometry 2 和自然语言推理工作由 Thang Luong 领导。
AlphaProof development was led by Thomas Hubert, Rishi Mehta and Laurent Sartran; AlphaGeometry 2 and natural language reasoning efforts were led by Thang Luong.
AlphaProof 的开发得到了 Hussain Masoom、Aja Huang、Miklós Z. Horváth、Tom Zahavy、Vivek Veeriah、Eric Wieser、Jessica Yung、Lei Yu、Yannick Schroecker、Julian Schrittwieser、Ottavia Bertolli、Borja Ibarz、Edward Lockhart、Edward Hughes、Mark Rowland、Grace Margand 的关键贡献。Alex Davies 和 Daniel Zheng 领导了非正式系统(如最终答案确定)的开发,关键贡献者包括 Iuliya Beloshapka、Ingrid von Glehn、Yin Li、Fabian Pedregosa、Ameya Velingker 和 Goran Žužić。Oliver Nash、Bhavik Mehta、Paul Lezeau、Salvatore Mercuri、Lawrence Wu、Calle Soenne、Thomas Murrills、Luigi Massacci 和 Andrew Yang 作为 Lean 专家提供建议和贡献。过去的贡献者包括 Amol Mandhane、Tom Eccles、Eser Aygün、Zhitao Gong、Richard Evans、Soňa Mokrá、Amin Barekatain、Wendy Shang、Hannah Openshaw、Felix Gimeno。这项工作由 David Silver 和 Pushmeet Kohli 指导。
AlphaProof was developed with key contributions from Hussain Masoom, Aja Huang, Miklós Z. Horváth, Tom Zahavy, Vivek Veeriah, Eric Wieser, Jessica Yung, Lei Yu, Yannick Schroecker, Julian Schrittwieser, Ottavia Bertolli, Borja Ibarz, Edward Lockhart, Edward Hughes, Mark Rowland, Grace Margand. Alex Davies and Daniel Zheng led the development of informal systems such as final answer determination, with key contributions from Iuliya Beloshapka, Ingrid von Glehn, Yin Li, Fabian Pedregosa, Ameya Velingker and Goran Žužić. Oliver Nash, Bhavik Mehta, Paul Lezeau, Salvatore Mercuri, Lawrence Wu, Calle Soenne, Thomas Murrills, Luigi Massacci and Andrew Yang advised and contributed as Lean experts. Past contributors include Amol Mandhane, Tom Eccles, Eser Aygün, Zhitao Gong, Richard Evans, Soňa Mokrá, Amin Barekatain, Wendy Shang, Hannah Openshaw, Felix Gimeno. This work was advised by David Silver and Pushmeet Kohli.
AlphaGeometry 2 的开发由 Trieu Trinh 和 Yuri Chervonyi 领导,关键贡献者包括 Mirek Olšák、Xiaomeng Yang、Hoang Nguyen、Junehyuk Jung、Dawsen Hwang 和 Marcelo Menegali。自然语言推理系统的开发由 Golnaz Ghiasi、Garrett Bingham、YaGuang Li 领导,关键贡献者包括 Swaroop Mishra、Nigamaa Nayakanti、Sidharth Mudgal、Qijun Tan、Junehyuk Jung、Hoang Nguyen、Alex Zhai、Dawsen Hwang、Mingyang Deng、Clara Huiyi Hu、Jarrod Kahn、Maciej Kula、Cosmo Du。AlphaGeometry 和自然语言推理系统均由 Quoc Le 指导。
The development of AlphaGeometry 2 was led by Trieu Trinh and Yuri Chervonyi, with key contributions by Mirek Olšák, Xiaomeng Yang, Hoang Nguyen, Junehyuk Jung, Dawsen Hwang and Marcelo Menegali. The development of the natural language reasoning system was led by Golnaz Ghiasi, Garrett Bingham, YaGuang Li, with key contributions by Swaroop Mishra, Nigamaa Nayakanti, Sidharth Mudgal, Qijun Tan, Junehyuk Jung, Hoang Nguyen, Alex Zhai, Dawsen Hwang, Mingyang Deng, Clara Huiyi Hu, Jarrod Kahn, Maciej Kula, Cosmo Du. Both AlphaGeometry and natural language reasoning systems were advised by Quoc Le.
David Silver、Quoc Le、Demis Hassabis 和 Pushmeet Kohli 协调并管理了整个项目。
David Silver, Quoc Le, Demis Hassabis, and Pushmeet Kohli coordinated and managed the overall project.
我们还要感谢 Insuk Seo、Evan Chen、Zigmars Rasscevskis、Kari Ragnarsson、Junhwi Bae、Jeonghyun Ahn、Jimin Kim、Hung Pham、Nguyen Nguyen、Son Pham 和 Pasin Manurangsi,他们帮助评估了我们语言推理系统的质量。感谢 Jeff Stanway、Jessica Lo、Erica Moreira、Petko Yotov 和 Kareem Ayoub 对算力提供和管理的支持。感谢 IMO 委员会的 Gregor Dolinar 教授和 Geoff Smith MBE 博士的支持与合作;以及 Tu Vu、Hanzhao Lin、Chenkai Kuang、Vikas Verma、Yifeng Lu、Xinyun Chen、Denny Zhou、Vihan Jain、Henryk Michalewski、Xavier Garcia、Arjun Kar、Lampros Lamprou、Kaushal Patel、Kelvin Xu、Ilya Tolstikhin、Olivier Bousquet、Anton Tsitsulin、Dustin Zelle、CJ Carey、Sam Blackwell、Abhi Rao、Vahab Mirrokni、Behnam Neyshabur、Ethan Dyer、Keith Rush、Moritz Firsching、Dan Shved、Ihar Bury、Divyanshu Ranjan、Hadi Hashemi、Alexei Bendebury、Soheil Hassas Yeganeh、Shibl Mourad、Simon Schmitt、Satinder Baveja、Chris Dyer、Jacob Austin、Wenda Li、Heng-tze Cheng、Ed Chi、Koray Kavukcuoglu、Oriol Vinyals、Jeff Dean 和 Sergey Brin 的支持和建议。
We’d also like to thank Insuk Seo, Evan Chen, Zigmars Rasscevskis, Kari Ragnarsson, Junhwi Bae, Jeonghyun Ahn, Jimin Kim, Hung Pham, Nguyen Nguyen, Son Pham, and Pasin Manurangsi who helped evaluate the quality of our language reasoning system. Jeff Stanway, Jessica Lo, Erica Moreira, Petko Yotov and Kareem Ayoub for their support for compute provision and management. Prof Gregor Dolinar and Dr Geoff Smith MBE from the IMO Board, for the support and collaboration; and Tu Vu, Hanzhao Lin, Chenkai Kuang, Vikas Verma, Yifeng Lu, Xinyun Chen, Denny Zhou, Vihan Jain, Henryk Michalewski, Xavier Garcia, Arjun Kar, Lampros Lamprou, Kaushal Patel, Kelvin Xu, Ilya Tolstikhin, Olivier Bousquet, Anton Tsitsulin, Dustin Zelle, CJ Carey, Sam Blackwell, Abhi Rao, Vahab Mirrokni, Behnam Neyshabur, Ethan Dyer, Keith Rush, Moritz Firsching, Dan Shved, Ihar Bury, Divyanshu Ranjan, Hadi Hashemi, Alexei Bendebury, Soheil Hassas Yeganeh, Shibl Mourad, Simon Schmitt, Satinder Baveja, Chris Dyer, Jacob Austin, Wenda Li, Heng-tze Cheng, Ed Chi, Koray Kavukcuoglu, Oriol Vinyals, Jeff Dean and Sergey Brin for their support and advice.
最后,我们要感谢 Lean 和 Mathlib 项目的众多贡献者,没有他们,AlphaProof 就不可能实现。
Finally, we’d like to thank the many contributors to the Lean and Mathlib projects, without whom AlphaProof wouldn’t have been possible.