BLIP3-o:全开放统一多模态模型家族——架构、训练与数据集

BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset

李俊男 Junnan Li · Salesforce · 2025-05-14 · arXiv:2505.09568 ↗ · 被引 322

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

统一图像理解与生成在近期多模态模型研究中日益受到关注。尽管图像理解的设计选择已被广泛研究,但针对包含图像生成的统一框架,其最优模型架构和训练方案仍待探索。受自回归模型和扩散模型在高质量生成与可扩展性方面强大潜力的启发,我们对其在统一多模态设置中的应用进行了全面研究,重点关注图像表示、建模目标和训练策略。基于这些研究,我们提出了一种新方法,采用扩散 Transformer 生成语义丰富的 CLIP 图像特征,而非传统的基于 VAE 的表示。该设计同时提高了训练效率和生成质量。此外,我们证明,统一模型的顺序预训练策略——先训练图像理解,再训练图像生成——通过保持图像理解能力同时发展强大的图像生成能力,提供了实际优势。最后,我们精心整理了一个高质量指令微调数据集 BLIP3o-60k,用于图像生成,通过使用涵盖各种场景、物体、人类手势等的多样化描述提示 GPT-4o 生成。基于创新的模型设计、训练方案和数据集,我们开发了 BLIP3-o,一套最先进的统一多模态模型。BLIP3-o 在大多数流行的图像理解和生成任务基准上均取得了优越性能。为促进未来研究,我们完全开源了模型,包括代码、模型权重、训练脚本以及预训练和指令微调数据集。

Unifying image understanding and generation has gained growing attention in recent research on multimodal models. Although design choices for image understanding have been extensively studied, the optimal model architecture and training recipe for a unified framework with image generation remain underexplored. Motivated by the strong potential of autoregressive and diffusion models for high-quality generation and scalability, we conduct a comprehensive study of their use in unified multimodal settings, with emphasis on image representations, modeling objectives, and training strategies. Grounded in these investigations, we introduce a novel approach that employs a diffusion transformer to generate semantically rich CLIP image features, in contrast to conventional VAE-based representations. This design yields both higher training efficiency and improved generative quality. Furthermore, we demonstrate that a sequential pretraining strategy for unified models-first training on image understanding and subsequently on image generation-offers practical advantages by preserving image understanding capability while developing strong image generation ability. Finally, we carefully curate a high-quality instruction-tuning dataset BLIP3o-60k for image generation by prompting GPT-4o with a diverse set of captions covering various scenes, objects, human gestures, and more. Building on our innovative model design, training recipe, and datasets, we develop BLIP3-o, a suite of state-of-the-art unified multimodal models. BLIP3-o achieves superior performance across most of the popular benchmarks spanning both image understanding and generation tasks. To facilitate future research, we fully open-source our models, including code, model weights, training scripts, and pretraining and instruction tuning datasets.

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 20)

阅读逐段中英对照全文 →