AlphaFold: a solution to a 50-year-old grand challenge in biology — Google DeepMind
打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→2022 年 7 月,我们发布了几乎所有已知科学目录中蛋白质的 AlphaFold 蛋白质结构预测。在此阅读最新博客。 蛋白质对生命至关重要,支撑着几乎所有的生命功能。它们是由氨基酸链组成的大型复杂分子,蛋白质的功能很大程度上取决于其独特的三维结构。弄清楚蛋白质折叠成什么形状被称为“蛋白质折叠问题”,并在过去 50 年中一直是生物学的一个重大挑战。在一项重大科学进展中,我们最新版本的 AI 系统 AlphaFold 被两年一度的蛋白质结构预测关键评估(CASP)组织者认可为这一重大挑战的解决方案。这一突破展示了 AI 对科学发现的影响及其在解释和塑造我们世界的一些最基础领域中显著加速进步的潜力。 蛋白质的形状与其功能密切相关,预测这种结构的能力有助于更深入地理解其功能和作用方式。许多世界最大的挑战,如开发疾病治疗方法或寻找分解工业废物的酶,从根本上都与蛋白质及其作用相关。
In July 2022, we released AlphaFold protein structure predictions for nearly all catalogued proteins known to science. Read the latest blog here. Proteins are essential to life, supporting practically all its functions. They are large complex molecules, made up of chains of amino acids, and what a protein does largely depends on its unique 3D structure. Figuring out what shapes proteins fold into is known as the “protein folding problem”, and has stood as a grand challenge in biology for the past 50 years. In a major scientific advance, the latest version of our AI system AlphaFold has been recognised as a solution to this grand challenge by the organisers of the biennial Critical Assessment of protein Structure Prediction (CASP).
2022 年 7 月,我们发布了几乎所有已知科学目录中蛋白质的 AlphaFold 蛋白质结构预测。在此阅读最新博客。
In July 2022, we released AlphaFold protein structure predictions for nearly all catalogued proteins known to science. Read the latest blog here.
蛋白质对生命至关重要,支撑着几乎所有的生命功能。它们是由氨基酸链组成的大型复杂分子,蛋白质的功能很大程度上取决于其独特的三维结构。弄清楚蛋白质折叠成什么形状被称为“蛋白质折叠问题”,并在过去 50 年中一直是生物学的一个重大挑战。在一项重大科学进展中,我们最新版本的 AI 系统 AlphaFold 被两年一度的蛋白质结构预测关键评估(CASP)组织者认可为这一重大挑战的解决方案。这一突破展示了 AI 对科学发现的影响及其在解释和塑造我们世界的一些最基础领域中显著加速进步的潜力。
Proteins are essential to life, supporting practically all its functions. They are large complex molecules, made up of chains of amino acids, and what a protein does largely depends on its unique 3D structure. Figuring out what shapes proteins fold into is known as the “protein folding problem”, and has stood as a grand challenge in biology for the past 50 years. In a major scientific advance, the latest version of our AI system AlphaFold has been recognised as a solution to this grand challenge by the organisers of the biennial Critical Assessment of protein Structure Prediction (CASP). This breakthrough demonstrates the impact AI can have on scientific discovery and its potential to dramatically accelerate progress in some of the most fundamental fields that explain and shape our world.
蛋白质的形状与其功能密切相关,预测这种结构的能力有助于更深入地理解其功能和作用方式。许多世界最大的挑战,如开发疾病治疗方法或寻找分解工业废物的酶,从根本上都与蛋白质及其作用相关。
A protein’s shape is closely linked with its function, and the ability to predict this structure unlocks a greater understanding of what it does and how it works. Many of the world’s greatest challenges, like developing treatments for diseases or finding enzymes that break down industrial waste, are fundamentally tied to proteins and the role they play.
CASP 联合创始人兼主席,马里兰大学
Co-founder and Chair of CASP, University of Maryland
多年来,这一直是密集科学研究的焦点,使用各种实验技术来检查和确定蛋白质结构,如核磁共振和 X 射线晶体学。这些技术以及像冷冻电镜这样的新方法,依赖于大量的试错,每个结构可能需要数年的艰苦和费力的工作,并且需要使用价值数百万美元的专用设备。
This has been a focus of intensive scientific research for many years, using a variety of experimental techniques to examine and determine protein structures, such as nuclear magnetic resonance and X-ray crystallography. These techniques, as well as newer methods like cryo-electron microscopy, depend on extensive trial and error, which can take years of painstaking and laborious work per structure, and require the use of multi-million dollar specialised equipment.
在 1972 年诺贝尔化学奖的获奖感言中,克里斯蒂安·安芬森著名地提出,理论上,蛋白质的氨基酸序列应完全决定其结构。这一假设引发了一场长达五十年的探索,旨在能够仅基于蛋白质的一维氨基酸序列通过计算预测其三维结构,作为这些昂贵且耗时的实验方法的补充替代方案。然而,一个主要挑战是,蛋白质在最终形成其三维结构之前理论上可能折叠的方式数量是天文数字。1969 年,西勒斯·莱文塔尔指出,通过暴力计算枚举一个典型蛋白质的所有可能构型所需的时间将超过已知宇宙的年龄——莱文塔尔估计一个典型蛋白质有 10^300 种可能的构象。然而在自然界中,蛋白质自发折叠,有些甚至在毫秒内完成——这种二分法有时被称为莱文塔尔悖论。
In his acceptance speech for the 1972 Nobel Prize in Chemistry, Christian Anfinsen famously postulated that, in theory, a protein’s amino acid sequence should fully determine its structure. This hypothesis sparked a five decade quest to be able to computationally predict a protein’s 3D structure based solely on its 1D amino acid sequence as a complementary alternative to these expensive and time consuming experimental methods. A major challenge, however, is that the number of ways a protein could theoretically fold before settling into its final 3D structure is astronomical. In 1969 Cyrus Levinthal noted that it would take longer than the age of the known universe to enumerate all possible configurations of a typical protein by brute force calculation – Levinthal estimated 10^300 possible conformations for a typical protein. Yet in nature, proteins fold spontaneously, some within milliseconds – a dichotomy sometimes referred to as Levinthal’s paradox.
1994 年,John Moult 教授和 Krzysztof Fidelis 教授创立了 CASP,作为每两年一次的盲测评估,旨在促进研究、监测进展并确立蛋白质结构预测的最新技术水平。它既是评估预测技术的黄金标准,也是一个基于共同事业的独特全球社区。关键的是,CASP 选择那些刚刚通过实验确定(有些在评估时仍在等待确定)的蛋白质结构作为目标,供团队测试其结构预测方法;这些结构不会提前公布。参与者必须盲目预测蛋白质的结构,随后将这些预测与可获得的实验真实数据进行对比。我们感谢 CASP 的组织者和整个社区,尤其是那些通过其结构实现这种严格评估的实验学家。
In 1994, Professor John Moult and Professor Krzysztof Fidelis founded CASP as a biennial blind assessment to catalyse research, monitor progress, and establish the state of the art in protein structure prediction. It is both the gold standard for assessing predictive techniques and a unique global community built on shared endeavour. Crucially, CASP chooses protein structures that have only very recently been experimentally determined (some were still awaiting determination at the time of the assessment) to be targets for teams to test their structure prediction methods against; they are not published in advance. Participants must blindly predict the structure of the proteins, and these predictions are subsequently compared to the ground truth experimental data when they become available. We’re indebted to CASP’s organisers and the whole community, not least the experimentalists whose structures enable this kind of rigorous assessment.
CASP 用于衡量预测准确性的主要指标是全局距离测试(GDT),其范围从 0 到 100。简单来说,GDT 可以近似理解为在正确位置阈值距离内的氨基酸残基(蛋白质链中的珠子)的百分比。根据 Moult 教授的说法,大约 90 GDT 的分数被非正式地认为与实验方法获得的结果相当。
The main metric used by CASP to measure the accuracy of predictions is the Global Distance Test (GDT) which ranges from 0-100. In simple terms, GDT can be approximately thought of as the percentage of amino acid residues (beads in the protein chain) within a threshold distance from the correct position. According to Professor Moult, a score of around 90 GDT is informally considered to be competitive with results obtained from experimental methods.
在今天发布的第 14 届 CASP 评估结果中,我们最新的 AlphaFold 系统在所有目标上的总体中位数得分为 92.4 GDT。这意味着我们的预测平均误差(RMSD)约为 1.6 埃,相当于一个原子的宽度(或 0.1 纳米)。即使对于最难的蛋白质目标,即最具挑战性的自由建模类别中的目标,AlphaFold 也达到了 87.0 GDT 的中位数得分(数据可在此处获取)。
In the results from the 14th CASP assessment, released today, our latest AlphaFold system achieves a median score of 92.4 GDT overall across all targets. This means that our predictions have an average error (RMSD) of approximately 1.6 Angstroms, which is comparable to the width of an atom (or 0.1 of a nanometer). Even for the very hardest protein targets, those in the most challenging free-modelling category, AlphaFold achieves a median score of 87.0 GDT (data available here).
在每次 CASP 中,自由建模类别中最佳团队的预测中位数准确性的改进,以最佳 5 个 GDT 衡量。
Improvements in the median accuracy of predictions in the free modelling category for the best team in each CASP, measured as best-of-5 GDT.
自由建模类别中两个蛋白质目标的示例。AlphaFold 预测的高度准确结构与实验结果对比。
Two examples of protein targets in the free modelling category. AlphaFold predicts highly accurate structures measured against experimental result.
这些令人兴奋的结果为生物学家使用计算结构预测作为科学研究的核心工具开辟了潜力。我们的方法可能对重要类别的蛋白质特别有用,例如膜蛋白,这些蛋白质很难结晶,因此难以通过实验确定。
These exciting results open up the potential for biologists to use computational structure prediction as a core tool in scientific research. Our methods may prove especially helpful for important classes of proteins, such as membrane proteins, that are very difficult to crystallise and therefore challenging to experimentally determine.
诺贝尔奖得主、英国皇家学会主席
Nobel Laureate and President of The Royal Society
我们首次参加 CASP13 是在 2018 年,当时我们推出了 AlphaFold 的初始版本,在参赛者中取得了最高准确率。随后,我们在《自然》杂志上发表了关于 CASP13 方法的论文,并附带了相关代码,这些工作启发了其他研究和社区开发的开源实现。现在,我们开发的新深度学习架构推动了 CASP14 方法的变革,使我们能够达到前所未有的准确度。这些方法借鉴了生物学、物理学和机器学习领域,当然也包括过去半个世纪中许多蛋白质折叠领域科学家的研究成果。
We first entered CASP13 in 2018 with our initial version of AlphaFold, which achieved the highest accuracy among participants. Afterwards, we published a paper on our CASP13 methods in Nature with associated code, which has gone on to inspire other work and community-developed open source implementations. Now, new deep learning architectures we’ve developed have driven changes in our methods for CASP14, enabling us to achieve unparalleled levels of accuracy. These methods draw inspiration from the fields of biology, physics, and machine learning, as well as of course the work of many scientists in the protein-folding field over the past half-century.
折叠的蛋白质可以被视为一个“空间图”,其中残基是节点,边连接空间上邻近的残基。这个图对于理解蛋白质内部的物理相互作用及其进化历史至关重要。对于用于 CASP14 的最新版 AlphaFold,我们创建了一个基于注意力机制的神经网络系统,进行端到端训练,该系统试图解释这个图的结构,同时对其构建的隐式图进行推理。它利用进化相关的序列、多序列比对(MSA)以及氨基酸残基对的表示来优化这个图。
A folded protein can be thought of as a “spatial graph”, where residues are the nodes and edges connect the residues in close proximity. This graph is important for understanding the physical interactions within proteins, as well as their evolutionary history. For the latest version of AlphaFold, used at CASP14, we created an attention-based neural network system, trained end-to-end, that attempts to interpret the structure of this graph, while reasoning over the implicit graph that it’s building. It uses evolutionarily related sequences, multiple sequence alignment (MSA), and a representation of amino acid residue pairs to refine this graph.
通过迭代这一过程,系统对蛋白质的底层物理结构形成了强有力的预测,并能在数天内确定高度准确的结构。此外,AlphaFold 可以利用内部置信度度量来预测每个预测蛋白质结构的哪些部分是可靠的。
By iterating this process, the system develops strong predictions of the underlying physical structure of the protein and is able to determine highly-accurate structures in a matter of days. Additionally, AlphaFold can predict which parts of each predicted protein structure are reliable using an internal confidence measure.
我们在公开数据上训练了这个系统,这些数据包括来自蛋白质数据银行的约 17 万个蛋白质结构,以及包含未知结构蛋白质序列的大型数据库。训练使用了大约 16 个 TPUv3(即 128 个 TPUv3 核心,大致相当于 100-200 个 GPU),运行数周,这在当今大多数大型先进机器学习模型中算是一个相对适中的算力投入。与我们的 CASP13 AlphaFold 系统一样,我们正在准备一篇关于该系统的论文,将在适当时候提交给同行评审期刊。
We trained this system on publicly available data consisting of ~170,000 protein structures from the protein data bank together with large databases containing protein sequences of unknown structure. It uses approximately 16 TPUv3s (which is 128 TPUv3 cores or roughly equivalent to ~100-200 GPUs) run over a few weeks, a relatively modest amount of compute in the context of most large state-of-the-art models used in machine learning today. As with our CASP13 AlphaFold system, we are preparing a paper on our system to submit to a peer-reviewed journal in due course.
主要神经网络模型架构的概述。该模型对进化相关的蛋白质序列以及氨基酸残基对进行操作,在两个表示之间迭代传递信息以生成结构。
An overview of the main neural network model architecture. The model operates over evolutionarily related protein sequences as well as amino acid residue pairs, iteratively passing information between both representations to generate a structure.
当 DeepMind 十年前成立时,我们希望有一天人工智能的突破能够作为一个平台,帮助我们加深对基础科学问题的理解。如今,经过四年的努力构建 AlphaFold,我们开始看到这一愿景成为现实,对药物设计和环境可持续性等领域产生了影响。
When DeepMind started a decade ago, we hoped that one day AI breakthroughs would help serve as a platform to advance our understanding of fundamental scientific problems. Now, after 4 years of effort building AlphaFold, we’re starting to see that vision realised, with implications for areas like drug design and environmental sustainability.
马克斯·普朗克发育生物学研究所所长、CASP 评估员 Andrei Lupas 教授告诉我们:“AlphaFold 惊人准确的模型使我们能够解决一个困扰我们近十年的蛋白质结构,重新启动了我们对信号如何跨细胞膜传递的研究。”
Professor Andrei Lupas, Director of the Max Planck Institute for Developmental Biology and a CASP assessor, let us know that, “AlphaFold’s astonishingly accurate models have allowed us to solve a protein structure we were stuck on for close to a decade, relaunching our effort to understand how signals are transmitted across cell membranes.”
我们对 AlphaFold 在生物学研究和更广泛世界中的影响持乐观态度,并期待与其他人合作,在未来几年进一步了解其潜力。在撰写同行评审论文的同时,我们也在探索如何以可扩展的方式更广泛地提供该系统。
We’re optimistic about the impact AlphaFold can have on biological research and the wider world, and excited to collaborate with others to learn more about its potential in the years ahead. Alongside working on a peer-reviewed paper, we’re exploring how best to provide broader access to the system in a scalable way.
与此同时,我们也在研究蛋白质结构预测如何通过与少数专业团队合作,帮助我们理解特定疾病,例如通过识别功能失常的蛋白质并推理它们如何相互作用。这些见解可以促进更精确的药物开发工作,补充现有的实验方法,更快地找到有前景的治疗方案。
In the meantime, we’re also looking into how protein structure predictions could contribute to our understanding of specific diseases with a small number of specialist groups, for example by helping to identify proteins that have malfunctioned and to reason about how they interact. These insights could enable more precise work on drug development, complementing existing experimental methods to find promising treatments faster.
博士、Calico 创始人兼首席执行官、基因泰克前董事长兼首席执行官
PhD, Founder and CEO Calico, Former Chairman and CEO Genentech
我们还看到迹象表明,蛋白质结构预测作为科学界开发的众多工具之一,可能在未来大流行应对工作中发挥作用。今年早些时候,我们预测了 SARS-CoV-2 病毒的几种蛋白质结构,包括 ORF3a,其结构此前未知。在 CASP14 上,我们预测了另一种冠状病毒蛋白 ORF8 的结构。实验人员迅速而令人印象深刻的工作现已确认了 ORF3a 和 ORF8 的结构。尽管这些结构具有挑战性且相关序列很少,但与实验确定的结构相比,我们的两个预测都达到了很高的准确度。
We’ve also seen signs that protein structure prediction could be useful in future pandemic response efforts, as one of many tools developed by the scientific community. Earlier this year, we predicted several protein structures of the SARS-CoV-2 virus, including ORF3a, whose structures were previously unknown. At CASP14, we predicted the structure of another coronavirus protein, ORF8. Impressively quick work by experimentalists has now confirmed the structures of both ORF3a and ORF8. Despite their challenging nature and having very few related sequences, we achieved a high degree of accuracy on both of our predictions when compared to their experimentally determined structures.
除了加速对已知疾病的理解,我们还对这些技术在探索我们目前没有模型的数亿种蛋白质方面的潜力感到兴奋——这是一片未知生物学的广阔领域。由于 DNA 指定了构成蛋白质结构的氨基酸序列,基因组学革命使得从自然世界中大规模读取蛋白质序列成为可能——通用蛋白质数据库(UniProt)中已有 1.8 亿个蛋白质序列,并且还在增加。相比之下,由于从序列到结构所需的实验工作,蛋白质数据库(PDB)中只有大约 17 万个蛋白质结构。在未确定的蛋白质中,可能有一些具有新颖且令人兴奋的功能——就像望远镜帮助我们更深入地观察未知宇宙一样,像 AlphaFold 这样的技术可能帮助我们找到它们。
As well as accelerating understanding of known diseases, we’re excited about the potential for these techniques to explore the hundreds of millions of proteins we don’t currently have models for – a vast terrain of unknown biology. Since DNA specifies the amino acid sequences that comprise protein structures, the genomics revolution has made it possible to read protein sequences from the natural world at massive scale – with 180 million protein sequences and counting in the Universal Protein database (UniProt). In contrast, given the experimental work needed to go from sequence to structure, only around 170,000 protein structures are in the Protein Data Bank (PDB). Among the undetermined proteins may be some with new and exciting functions and – just as a telescope helps us see deeper into the unknown universe – techniques like AlphaFold may help us find them.
AlphaFold 是我们迄今为止最重要的进展之一,但与所有科学研究一样,仍有许多问题有待解答。并非我们预测的每个结构都是完美的。还有很多需要学习的地方,包括多个蛋白质如何形成复合物,它们如何与 DNA、RNA 或小分子相互作用,以及我们如何确定所有氨基酸侧链的精确位置。与他人合作,我们还需要学习如何最好地利用这些科学发现来开发新药物、管理环境等。
AlphaFold is one of our most significant advances to date but, as with all scientific research, there are still many questions to answer. Not every structure we predict will be perfect. There’s still much to learn, including how multiple proteins form complexes, how they interact with DNA, RNA, or small molecules, and how we can determine the precise location of all amino acid side chains. In collaboration with others, there’s also much to learn about how best to use these scientific discoveries in the development of new medicines, ways to manage the environment, and more.
对于所有从事科学计算和机器学习方法的人来说,像 AlphaFold 这样的系统展示了人工智能作为帮助基础发现工具的惊人潜力。正如 50 年前 Anfinsen 提出了一个当时远远超出科学能力范围的挑战一样,我们宇宙的许多方面仍然未知。今天宣布的进展让我们更加相信,人工智能将成为人类扩展科学知识前沿的最有用工具之一,我们期待着未来多年的辛勤工作和发现!
For all of us working on computational and machine learning methods in science, systems like AlphaFold demonstrate the stunning potential for AI as a tool to aid fundamental discovery. Just as 50 years ago Anfinsen laid out a challenge far beyond science’s reach at the time, there are many aspects of our universe that remain unknown. The progress announced today gives us further confidence that AI will become one of humanity’s most useful tools in expanding the frontiers of scientific knowledge, and we’re looking forward to the many years of hard work and discovery ahead!
在我们发表关于这项工作的论文之前,请引用:
Until we’ve published a paper on this work, please cite:
使用深度学习进行高精度蛋白质结构预测
High Accuracy Protein Structure Prediction Using Deep Learning
John Jumper, Richard Evans, Alexander Pritzel, Tim Green, Michael Figurnov, Kathryn Tunyasuvunakool, Olaf Ronneberger, Russ Bates, Augustin Žídek, Alex Bridgland, Clemens Meyer, Simon A A Kohl, Anna Potapenko, Andrew J Ballard, Andrew Cowie, Bernardino Romera-Paredes, Stanislav Nikolov, Rishub Jain, Jonas Adler, Trevor Back, Stig Petersen, David Reiman, Martin Steinegger, Michalina Pacholska, David Silver, Oriol Vinyals, Andrew W Senior, Koray Kavukcuoglu, Pushmeet Kohli, Demis Hassabis.
John Jumper, Richard Evans, Alexander Pritzel, Tim Green, Michael Figurnov, Kathryn Tunyasuvunakool, Olaf Ronneberger, Russ Bates, Augustin Žídek, Alex Bridgland, Clemens Meyer, Simon A A Kohl, Anna Potapenko, Andrew J Ballard, Andrew Cowie, Bernardino Romera-Paredes, Stanislav Nikolov, Rishub Jain, Jonas Adler, Trevor Back, Stig Petersen, David Reiman, Martin Steinegger, Michalina Pacholska, David Silver, Oriol Vinyals, Andrew W Senior, Koray Kavukcuoglu, Pushmeet Kohli, Demis Hassabis.
在第十四届蛋白质结构预测关键技术评估(摘要书)中,2020 年 11 月 30 日至 12 月 4 日。检索自此处。
In Fourteenth Critical Assessment of Techniques for Protein Structure Prediction (Abstract Book), 30 November - 4 December 2020. Retrieved from here.