自然语言处理中基于大量数据预训练模型的近期突破,为计算机视觉领域类似的基础模型开辟了道路。这类模型通过生成通用视觉特征(即无需微调即可跨图像分布和任务工作的特征),能极大简化图像在任何系统中的使用。本研究表明,现有的预训练方法,特别是自监督方法,如果在来自不同来源的足够精选数据上进行训练,就能产生此类特征。我们重新审视现有方法,并结合不同技术来扩展数据和模型规模的预训练。大部分技术贡献旨在加速和稳定大规模训练。在数据方面,我们提出了一种自动流水线,用于构建专用、多样且精选的图像数据集,而非自监督文献中通常使用的未精选数据。在模型方面,我们训练了一个具有 10 亿参数的 ViT 模型(Dosovitskiy 等人,2020),并将其蒸馏成一系列较小的模型,这些模型在图像和像素级别的大多数基准测试中超越了最佳可用的通用特征 OpenCLIP(Ilharco 等人,2021)。
The recent breakthroughs in natural language processing for model pretraining on large quantities of data have opened the way for similar foundation models in computer vision. These models could greatly simplify the use of images in any system by producing all-purpose visual features, i.e., features that work across image distributions and tasks without finetuning. This work shows that existing pretraining methods, especially self-supervised methods, can produce such features if trained on enough curated data from diverse sources. We revisit existing approaches and combine different techniques to scale our pretraining in terms of data and model size. Most of the technical contributions aim at accelerating and stabilizing the training at scale. In terms of data, we propose an automatic pipeline to build a dedicated, diverse, and curated image dataset instead of uncurated data, as typically done in the self-supervised literature. In terms of models, we train a ViT model (Dosovitskiy et al., 2020) with 1B parameters and distill it into a series of smaller models that surpass the best available all-purpose features, OpenCLIP (Ilharco et al., 2021) on most of the benchmarks at image and pixel levels.
核心贡献 · Key contributions
提出 DINOv2,一种从大规模精选数据中无监督学习鲁棒视觉特征的自监督方法。 Proposes DINOv2, a self-supervised method for learning robust visual features from large curated data without supervision.
引入自动流水线,从非精选网络数据构建多样化的精选图像数据集(LVD-142M)。 Introduces an automatic pipeline to build a diverse and curated image dataset (LVD-142M) from uncurated web data.
在图像级和像素级任务上使用冻结特征达到最先进性能,匹配或超越弱监督方法。 Achieves state-of-the-art performance on image-level and pixel-level tasks with frozen features, matching or surpassing weakly-supervised methods.
证明扩展数据和模型规模可提升特征质量,蒸馏使小模型受益于大模型。 Demonstrates that scaling data and model size improves feature quality, with distillation enabling smaller models to benefit from larger ones.
提供大规模稳定高效训练的技术改进,包括更快的训练和更低的内存使用。 Provides technical improvements for stable and efficient training at scale, including faster training and reduced memory usage.
展示涌现特性如物体部件理解和场景几何,具有进一步扩展的潜力。 Shows emergent properties like object part understanding and scene geometry, with potential for further scaling.
局限 · Limitations
该方法需要大规模精选数据集和大量计算资源进行训练。 The method requires large curated datasets and significant computational resources for training.
在 Places205 等基准上性能落后于弱监督模型。 Performance on some benchmarks like Places205 lags behind weakly-supervised models.
地理和收入偏差仍然存在,在非洲地区和低收入家庭上性能较低。 Geographical and income biases persist, with lower performance on African regions and low-income households.
研究未充分探索微调超参数,未挖掘潜在增益。 The study does not fully explore finetuning hyperparameters, leaving potential gains unexplored.
碳足迹显著,项目总排放量估计在 0.5 千至 1 千吨二氧化碳当量之间。 Carbon footprint is substantial, with total project emissions estimated between 0.5k and 1k tCO2eq.