几何深度学习:网格、群、图、测地线与规范

Geometric Deep Learning: Grids, Groups, Graphs, Geodesics, and Gauges

佩塔尔·韦利奇科维奇 Petar Veličković · · 2021-04-27 · arXiv:2104.13478 ↗ · 被引 1682

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

过去十年见证了数据科学和机器学习中的一场实验革命,其典型代表便是深度学习方法。事实上,许多先前被认为遥不可及的高维学习任务——如计算机视觉、下围棋或蛋白质折叠——在适当的计算规模下确实可行。值得注意的是,深度学习的本质建立在两个简单的算法原则之上:其一,表示学习或特征学习的概念,即通过自适应的、往往是层次化的特征来捕捉每项任务相应的规律性;其二,通过局部梯度下降类方法进行学习,通常以反向传播实现。虽然在高维空间中学习通用函数是一个受维数灾难困扰的估计问题,但大多数感兴趣的任务并非通用,它们带有由物理世界底层低维性和结构所决定的基本预定义规律。本文旨在通过统一的几何原理来揭示这些规律,这些原理可广泛应用于各类领域。这种秉承菲利克斯·克莱因埃朗根纲领精神的“几何统一”努力具有双重目的:一方面,它为研究最成功的神经网络架构(如 CNN、RNN、GNN 和 Transformer)提供了一个共同的数学框架;另一方面,它提供了一种建设性程序,将先验物理知识融入神经架构,并为尚未发明的未来架构提供原则性的构建方法。

The last decade has witnessed an experimental revolution in data science and machine learning, epitomised by deep learning methods. Indeed, many high-dimensional learning tasks previously thought to be beyond reach -- such as computer vision, playing Go, or protein folding -- are in fact feasible with appropriate computational scale. Remarkably, the essence of deep learning is built from two simple algorithmic principles: first, the notion of representation or feature learning, whereby adapted, often hierarchical, features capture the appropriate notion of regularity for each task, and second, learning by local gradient-descent type methods, typically implemented as backpropagation. While learning generic functions in high dimensions is a cursed estimation problem, most tasks of interest are not generic, and come with essential pre-defined regularities arising from the underlying low-dimensionality and structure of the physical world. This text is concerned with exposing these regularities through unified geometric principles that can be applied throughout a wide spectrum of applications. Such a 'geometric unification' endeavour, in the spirit of Felix Klein's Erlangen Program, serves a dual purpose: on one hand, it provides a common mathematical framework to study the most successful neural network architectures, such as CNNs, RNNs, GNNs, and Transformers. On the other hand, it gives a constructive procedure to incorporate prior physical knowledge into neural architectures and provide principled way to build future architectures yet to be invented.

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 46)

阅读逐段中英对照全文 →