AlphaEvolve:由 Gemini 驱动的编码代理,用于设计高级算法——Google DeepMind

AlphaEvolve: A Gemini-powered coding agent for designing advanced algorithms — Google DeepMind

谷歌 DeepMind Google DeepMind · Google DeepMind · 2025-05-14 · DeepMind Blog ↗

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

新的人工智能智能体通过将大型语言模型的创造力与自动评估器相结合,为数学和计算的实际应用演化算法。 大型语言模型(LLM)具有显著的通用性。它们可以总结文档、生成代码,甚至构思新想法。现在,我们已将这些能力扩展到针对数学和现代计算中基础且高度复杂的问题。 今天,我们宣布推出 AlphaEvolve,这是一个由大型语言模型驱动的演化编码智能体,用于通用算法的发现和优化。AlphaEvolve 将我们 Gemini 模型的创造性问题解决能力与验证答案的自动评估器相结合,并利用演化框架来改进最有前景的想法。

New AI agent evolves algorithms for math and practical applications in computing by combining the creativity of large language models with automated evaluators Large language models (LLMs) are remarkably versatile. They can summarize documents, generate code or even brainstorm new ideas. And now we’ve expanded these capabilities to target fundamental and highly complex problems in mathematics and modern computing. Today, we’re announcing AlphaEvolve, an evolutionary coding agent powered by large language models for general-purpose algorithm discovery and optimization. AlphaEvolve pairs the creative problem-solving capabilities of our Gemini models with automated evaluators that verify answers, and uses an evolutionary framework to improve upon the most promising ideas.

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 8)

全文 · Full text(逐段中英对照)

AlphaEvolve:一个由 Gemini 驱动的编码智能体,用于设计高级算法 AlphaEvolve: A Gemini-powered coding agent for designing advanced algorithms

新的人工智能智能体通过将大型语言模型的创造力与自动评估器相结合,为数学和计算的实际应用演化算法。

New AI agent evolves algorithms for math and practical applications in computing by combining the creativity of large language models with automated evaluators

大型语言模型(LLM)具有显著的通用性。它们可以总结文档、生成代码,甚至构思新想法。现在,我们已将这些能力扩展到针对数学和现代计算中基础且高度复杂的问题。

Large language models (LLMs) are remarkably versatile. They can summarize documents, generate code or even brainstorm new ideas. And now we’ve expanded these capabilities to target fundamental and highly complex problems in mathematics and modern computing.

今天,我们宣布推出 AlphaEvolve,这是一个由大型语言模型驱动的演化编码智能体,用于通用算法的发现和优化。AlphaEvolve 将我们 Gemini 模型的创造性问题解决能力与验证答案的自动评估器相结合,并利用演化框架来改进最有前景的想法。

Today, we’re announcing AlphaEvolve, an evolutionary coding agent powered by large language models for general-purpose algorithm discovery and optimization. AlphaEvolve pairs the creative problem-solving capabilities of our Gemini models with automated evaluators that verify answers, and uses an evolutionary framework to improve upon the most promising ideas.

AlphaEvolve 提升了谷歌数据中心、芯片设计和 AI 训练流程的效率——包括训练 AlphaEvolve 自身所依赖的大型语言模型。它还帮助设计了更快的矩阵乘法算法,并为开放的数学问题找到了新的解决方案,显示出在众多领域应用的巨大潜力。

AlphaEvolve enhanced the efficiency of Google's data centers, chip design and AI training processes — including training the large language models underlying AlphaEvolve itself. It has also helped design faster matrix multiplication algorithms and find new solutions to open mathematical problems, showing incredible promise for application across many areas.

利用大语言模型设计更好的算法 Designing better algorithms with large language models

2023 年,我们首次证明大语言模型可以生成计算机代码形式的函数,以帮助在开放科学问题上发现新的、可证明正确的知识。AlphaEvolve 是一个智能体,它能够超越单一函数发现,进化整个代码库并开发更复杂的算法。

In 2023, we showed for the first time that large language models can generate functions written in computer code to help discover new and provably correct knowledge on an open scientific problem. AlphaEvolve is an agent that can go beyond single function discovery to evolve entire codebases and develop much more complex algorithms.

AlphaEvolve 利用一组最先进的大语言模型:我们最快、最高效的模型 Gemini Flash 最大化探索思路的广度,而我们最强大的模型 Gemini Pro 则通过富有洞察力的建议提供关键深度。这些模型共同提出实现算法解决方案的计算机程序代码。

AlphaEvolve leverages an ensemble of state-of-the-art large language models: our fastest and most efficient model, Gemini Flash, maximizes the breadth of ideas explored, while our most powerful model, Gemini Pro, provides critical depth with insightful suggestions. Together, these models propose computer programs that implement algorithmic solutions as code.

图示展示了提示采样器如何首先为大语言模型组装提示,然后模型生成新程序。这些程序由评估器评估并存储在程序数据库中。该数据库实现了一个进化算法,决定哪些程序将用于未来的提示。

Diagram showing how the prompt sampler first assembles a prompt for the language models, which then generate new programs. These programs are evaluated by evaluators and stored in the programs database. This database implements an evolutionary algorithm that determines which programs will be used for future prompts.

AlphaEvolve 使用自动评估指标验证、运行和评分所提出的程序。这些指标提供了对每个解决方案准确性和质量的客观、可量化的评估。这使得 AlphaEvolve 在数学和计算机科学等可以清晰、系统衡量进展的广泛领域中特别有用。

AlphaEvolve verifies, runs and scores the proposed programs using automated evaluation metrics. These metrics provide an objective, quantifiable assessment of each solution’s accuracy and quality. This makes AlphaEvolve particularly helpful in a broad range of domains where progress can be clearly and systematically measured, like in math and computer science.

优化我们的计算生态系统 Optimizing our computing ecosystem

过去一年,我们将 AlphaEvolve 发现的算法部署到谷歌的计算生态系统中,包括数据中心、硬件和软件。这些改进的每一项影响都在我们的人工智能和计算基础设施上被放大,从而为所有用户构建一个更强大、更可持续的数字生态系统。

Over the past year, we’ve deployed algorithms discovered by AlphaEvolve across Google’s computing ecosystem, including our data centers, hardware and software. The impact of each of these improvements is multiplied across our AI and computing infrastructure to build a more powerful and sustainable digital ecosystem for all our users.

图表展示了 AlphaEvolve 如何帮助谷歌提供更高效的数字生态系统,从数据中心调度和硬件设计到 AI 模型训练。

Diagram showing how AlphaEvolve helps Google deliver a more efficient digital ecosystem, from data center scheduling and hardware design to AI model training.

改进数据中心调度 Improving data center scheduling

AlphaEvolve 发现了一种简单但极其有效的启发式方法,帮助 Borg 更高效地编排 Google 庞大的数据中心。该解决方案已投入生产超过一年,平均持续回收 Google 全球算力的 0.7%。这一持续的效率提升意味着,在任何给定时刻,相同的计算资源可以完成更多任务。AlphaEvolve 的解决方案不仅带来了强劲的性能,还提供了人类可读代码的重要操作优势:可解释性、可调试性、可预测性和易于部署。

AlphaEvolve discovered a simple yet remarkably effective heuristic to help Borg orchestrate Google's vast data centers more efficiently. This solution, now in production for over a year, continuously recovers, on average, 0.7% of Google’s worldwide compute resources. This sustained efficiency gain means that at any given moment, more tasks can be completed on the same computational footprint. AlphaEvolve's solution not only leads to strong performance but also offers significant operational advantages of human-readable code: interpretability, debuggability, predictability and ease of deployment.

辅助硬件设计 Assisting in hardware design

AlphaEvolve 提出了一种 Verilog 重写方案,在用于矩阵乘法的关键、高度优化的算术电路中移除了不必要的位。至关重要的是,该方案必须通过稳健的验证方法,以确认修改后的电路保持功能正确性。该方案已被集成到即将推出的张量处理单元(TPU)中,这是谷歌定制的 AI 加速器。通过以芯片设计师的标准语言提出修改建议,AlphaEvolve 促进了 AI 与硬件工程师之间的协作,以加速未来专用芯片的设计。

AlphaEvolve proposed a Verilog rewrite that removed unnecessary bits in a key, highly optimized arithmetic circuit for matrix multiplication. Crucially, the proposal must pass robust verification methods to confirm that the modified circuit maintains functional correctness. This proposal was integrated into an upcoming Tensor Processing Unit (TPU), Google’s custom AI accelerator. By suggesting modifications in the standard language of chip designers, AlphaEvolve promotes a collaborative approach between AI and hardware engineers to accelerate the design of future specialized chips.

增强 AI 训练与推理 Enhancing AI training and inference

AlphaEvolve 正在加速 AI 性能和研发速度。通过寻找更智能的方法将大型矩阵乘法运算分解为更易管理的子问题,它将 Gemini 架构中这一关键内核的速度提升了 23%,从而使 Gemini 的训练时间减少了 1%。由于开发生成式 AI 模型需要大量的算力资源,每一点效率提升都意味着可观的节省。除了性能提升,AlphaEvolve 还显著减少了内核优化所需的工程时间,从数周的专家工作缩短为几天的自动化实验,使研究人员能够更快地创新。

AlphaEvolve is accelerating AI performance and research velocity. By finding smarter ways to divide a large matrix multiplication operation into more manageable subproblems, it sped up this vital kernel in Gemini’s architecture by 23%, leading to a 1% reduction in Gemini's training time. Because developing generative AI models requires substantial computing resources, every efficiency gained translates to considerable savings. Beyond performance gains, AlphaEvolve significantly reduces the engineering time required for kernel optimization, from weeks of expert effort to days of automated experiments, allowing researchers to innovate faster.

AlphaEvolve 还可以优化底层 GPU 指令。这一极其复杂的领域通常已被编译器高度优化,因此人类工程师通常不会直接修改它。AlphaEvolve 在基于 Transformer 的 AI 模型中实现了 FlashAttention 内核实现高达 32.5% 的加速。这种优化有助于专家定位性能瓶颈,并轻松地将改进整合到他们的代码库中,从而提高他们的生产力,并实现未来在算力和能源方面的节省。

AlphaEvolve can also optimize low level GPU instructions. This incredibly complex domain is usually already heavily optimized by compilers, so human engineers typically don't modify it directly. AlphaEvolve achieved up to a 32.5% speedup for theFlashAttention kernel implementation inTransformer-based AI models. This kind of optimization helps experts pinpoint performance bottlenecks and easily incorporate the improvements into their codebase, boosting their productivity and enabling future savings in compute and energy.

推进数学与算法发现的前沿 Advancing the frontiers in mathematics and algorithm discovery

AlphaEvolve 还能为复杂的数学问题提出新方法。给定一个计算机程序的最小代码框架,AlphaEvolve 设计了一种新颖的基于梯度的优化过程的多个组件,该过程发现了多个用于矩阵乘法(计算机科学中的一个基本问题)的新算法。

AlphaEvolve can also propose new approaches to complex mathematical problems. Provided with a minimal code skeleton for a computer program, AlphaEvolve designed many components of a novel gradient-based optimization procedure that discovered multiple new algorithms for matrix multiplication, a fundamental problem in computer science.

AlphaEvolve 为发现更快的矩阵乘法算法所提出的更改列表。在此示例中,AlphaEvolve 提出了跨多个组件的广泛更改,包括优化器和权重初始化、损失函数以及超参数扫描。这些更改非常不平凡,在进化过程中需要 15 次突变。

A list of changes proposed by AlphaEvolve to discover faster matrix multiplication algorithms. In this example, AlphaEvolve proposes extensive changes across several components, including the optimizer and weight initialization, the loss function, and hyperparameter sweep. These changes are highly non-trivial, requiring 15 mutations during the evolutionary process.

AlphaEvolve 的过程发现了一种使用 48 次标量乘法计算 4x4 复值矩阵的算法,改进了 Strassen 1969 年的算法,该算法此前被认为是该设置下的最佳算法。这一发现表明,与我们之前专门研究矩阵乘法算法的工作 AlphaTensor 相比,取得了显著进步;对于 4x4 矩阵,AlphaTensor 仅发现了二进制算术的改进。

AlphaEvolve’s procedure found an algorithm to multiply 4x4 complex-valued matrices using 48 scalar multiplications, improving upon Strassen’s 1969 algorithm that was previously known as the best in this setting. This finding demonstrates a significant advance over our previous work, AlphaTensor, which specialized in matrix multiplication algorithms, and for 4x4 matrices, only found improvements for binary arithmetic.

为了探究 AlphaEvolve 的广度,我们将该系统应用于数学分析、几何、组合学和数论中的 50 多个开放问题。该系统的灵活性使我们能够在几小时内设置大多数实验。据我们所知,在大约 75% 的情况下,它重新发现了最先进的解决方案。

To investigate AlphaEvolve’s breadth, we applied the system to over 50 open problems in mathematical analysis, geometry, combinatorics and number theory. The system’s flexibility enabled us to set up most experiments in a matter of hours. In roughly 75% of cases, it rediscovered state-of-the-art solutions, to the best of our knowledge.

在 20% 的情况下,AlphaEvolve 改进了先前已知的最佳解决方案,在相应的开放问题上取得了进展。例如,它推进了亲吻数问题。这个几何挑战已经让数学家着迷了 300 多年,涉及与一个公共单位球相切的最大非重叠球体数量。AlphaEvolve 发现了一个由 593 个外球组成的构型,并在 11 维中建立了新的下界。

And in 20% of cases, AlphaEvolve improved the previously best known solutions, making progress on the corresponding open problems. For example, it advanced the kissing number problem. This geometric challenge has fascinated mathematicians for over 300 years and concerns the maximum number of non-overlapping spheres that touch a common unit sphere. AlphaEvolve discovered a configuration of 593 outer spheres and established a new lower bound in 11 dimensions.

前进之路 The path forward

AlphaEvolve 展示了从发现特定领域的算法到为广泛现实挑战开发更复杂算法的演进过程。我们期待 AlphaEvolve 能随着大语言模型能力的提升而持续改进,尤其是在它们变得更擅长编程时。

AlphaEvolve displays the progression from discovering algorithms for specific domains to developing more complex algorithms for a wide range of real-world challenges. We’re expecting AlphaEvolve to continue improving alongside the capabilities of large language models, especially as they become even better at coding.

我们与 People + AI Research 团队合作,构建了一个友好的用户界面来与 AlphaEvolve 交互。我们计划为选定的学术用户提供早期访问计划,并正在探索让 AlphaEvolve 更广泛可用的可能性。要注册您的兴趣,请填写此表格。

Together with the People + AI Research team, we’ve been building a friendly user interface for interacting with AlphaEvolve. We’re planning an Early Access Program for selected academic users and also exploring possibilities to make AlphaEvolve more broadly available. To register your interest, please complete this form.

虽然 AlphaEvolve 目前应用于数学和计算领域,但其通用性意味着它可以应用于任何解决方案可描述为算法并能自动验证的问题。我们相信 AlphaEvolve 可能在更多领域带来变革,如材料科学、药物发现、可持续性以及更广泛的技术和商业应用。

While AlphaEvolve is currently being applied across math and computing, its general nature means it can be applied to any problem whose solution can be described as an algorithm, and automatically verified. We believe AlphaEvolve could be transformative across many more areas such as material science, drug discovery, sustainability and wider technological and business applications.

在我们的白皮书中阅读更多详情注册您对使用 AlphaEvolve 的兴趣在我们的 Google Colab 中查看 AlphaEvolve 的数学结果

Read more details in our white paperRegister your interest in using AlphaEvolveSee AlphaEvolve’s mathematical results in our Google Colab

AlphaEvolve 由 Matej Balog、Alexander Novikov、Ngân Vũ、Marvin Eisenberger、Emilien Dupont、Po-Sen Huang、Adam Zsolt Wagner、Sergey Shirobokov、Borislav Kozlovskii、Francisco J. R. Ruiz、Abbas Mehrabian、M. Pawan Kumar、Abigail See、Swarat Chaudhuri、George Holland、Alex Davies、Sebastian Nowozin 和 Pushmeet Kohli 开发。这项研究是我们专注于使用 AI 进行算法发现工作的一部分。

AlphaEvolve was developed by Matej Balog, Alexander Novikov, Ngân Vũ, Marvin Eisenberger, Emilien Dupont, Po-Sen Huang, Adam Zsolt Wagner, Sergey Shirobokov, Borislav Kozlovskii, Francisco J. R. Ruiz, Abbas Mehrabian, M. Pawan Kumar, Abigail See, Swarat Chaudhuri, George Holland, Alex Davies, Sebastian Nowozin, and Pushmeet Kohli. This research was developed as part of our effort focused on using AI for algorithm discovery.

我们衷心感谢 Jean-Baptiste Alayrac、Ankit Anand、Natasha Antropova、Giorgio Arena、Mohammadamin Barekatain、Johannes Bausch、Henning Becker、Daniel Belov、Alexander Belyaev、Sebastian Bodenstein、Sebastian Borgeaud、Calin Cascaval、Indranil Chakraborty、Benjamin Chetioui、Justin Chiu、Christopher Clark、Marco Cornero、Jeff Dean、Gaurav Dhiman、Yanislav Donchev、Srikanth Dwarakanath、Jordan Ellenberg、Alhussein Fawzi、Michael Figurnov、Aaron Gentleman、Bogdan Georgiev、Sergio Guadarrama、Demis Hassabis、Patrick Heisel、Chase Hensel、Koray Kavukcuoglu、Sultan Kenjeyev、Aliia Khasanova、Sridhar Lakshmanamurthy、Sergei Lebedev、Dmitry Lepikhin、Daniel Mankowitz、Andrea Michi、Kieran Milan、Vinod Nair、Robert O'Callahan、Cosmin Paduraru、Stig Petersen、Federico Piccinini、Parthasarathy Ranganatha、Bernardino Romera-Paredes、Georges Rotival、Kirk Sanders、Javier Gomez Serrano、Oleg Shyshkov、Timur Sitdikov、Tammo Spalink、Kerry Takenaka、Richard Tanburn、Terence Tao、Amin Vahdat、JD Velasquez、Dimitrios Vytiniotis、Julian Walker 和 Pengming Wang 的贡献、建议和支持。更多详情请参阅我们的白皮书。

We gratefully acknowledge contributions, advice, and support from Jean-Baptiste Alayrac, Ankit Anand, Natasha Antropova, Giorgio Arena, Mohammadamin Barekatain, Johannes Bausch, Henning Becker, Daniel Belov, Alexander Belyaev, Sebastian Bodenstein, Sebastian Borgeaud, Calin Cascaval, Indranil Chakraborty, Benjamin Chetioui, Justin Chiu, Christopher Clark, Marco Cornero, Jeff Dean, Gaurav Dhiman, Yanislav Donchev, Srikanth Dwarakanath, Jordan Ellenberg, Alhussein Fawzi, Michael Figurnov, Aaron Gentleman, Bogdan Georgiev, Sergio Guadarrama, Demis Hassabis, Patrick Heisel, Chase Hensel, Koray Kavukcuoglu, Sultan Kenjeyev, Aliia Khasanova, Sridhar Lakshmanamurthy, Sergei Lebedev, Dmitry Lepikhin, Daniel Mankowitz, Andrea Michi, Kieran Milan, Vinod Nair, Robert O'Callahan, Cosmin Paduraru, Stig Petersen, Federico Piccinini, Parthasarathy Ranganatha, Bernardino Romera-Paredes, Georges Rotival, Kirk Sanders, Javier Gomez Serrano, Oleg Shyshkov, Timur Sitdikov, Tammo Spalink, Kerry Takenaka, Richard Tanburn, Terence Tao, Amin Vahdat, JD Velasquez, Dimitrios Vytiniotis, Julian Walker, and Pengming Wang. For more details, please see our white paper.

我们感谢 Armin Senoner、Juanita Bawagan、Jane Park、Arielle Bier 和 Molly Beck 对博客文章的反馈以及对本公告的帮助;感谢 William Hood、Irina Andronic、Victoria Johnston、Lucas Dixon、Adam Connors 和 Jimbo Wilson 在插图和图表方面的帮助。

We would like to thank Armin Senoner, Juanita Bawagan, Jane Park, Arielle Bier, and Molly Beck for feedback on the blog post and help with this announcement; William Hood, Irina Andronic, Victoria Johnston, Lucas Dixon, Adam Connors, and Jimbo Wilson for help with the illustrations and figures

注册以获取我们最新创新的更新

Sign up for updates on our latest innovations

我接受 Google 的条款和条件,并确认我的信息将按照 Google 的隐私政策使用。

I accept Google's Terms and Conditions and acknowledge that my information will be used in accordance with Google's Privacy Policy.

Gemini 应用 Google AI StudioGoogle Antigravity

Gemini appGoogle AI StudioGoogle Antigravity

互动版:图/公式 + 针对本篇提问 →