统一多模态预训练中的涌现特性

Emerging Properties in Unified Multimodal Pretraining

邓超瑞 Chaorui Deng · ByteDance Seed · 2025-05-20 · arXiv:2505.14683 ↗ · 被引 727

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

统一多模态理解与生成在前沿专有系统中展现出令人印象深刻的能力。本文介绍 BAGEL,一个原生支持多模态理解与生成的开源基础模型。BAGEL 是一个统一的仅解码器模型,在来自大规模交错文本、图像、视频和网络数据的数万亿 token 上预训练。当使用如此多样化的多模态交错数据进行扩展时,BAGEL 在复杂多模态推理中展现出涌现能力。因此,它在标准基准测试中显著优于开源统一模型的多模态生成和理解,同时展现出高级多模态推理能力,如自由形式图像操作、未来帧预测、3D 操作和世界导航。为了促进多模态研究的更多机会,我们分享了关键发现、预训练细节、数据创建协议,并向社区发布了我们的代码和检查点。项目页面位于 https://bagel-ai.org/。

Unifying multimodal understanding and generation has shown impressive capabilities in cutting-edge proprietary systems. In this work, we introduce BAGEL, an open-source foundational model that natively supports multimodal understanding and generation. BAGEL is a unified, decoder-only model pretrained on trillions of tokens curated from large-scale interleaved text, image, video, and web data. When scaled with such diverse multimodal interleaved data, BAGEL exhibits emerging capabilities in complex multimodal reasoning. As a result, it significantly outperforms open-source unified models in both multimodal generation and understanding across standard benchmarks, while exhibiting advanced multimodal reasoning abilities such as free-form image manipulation, future frame prediction, 3D manipulation, and world navigation. In the hope of facilitating further opportunities for multimodal research, we share the key findings, pretraining details, data creation protocal, and release our code and checkpoints to the community. The project page is at https://bagel-ai.org/

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 25)

阅读逐段中英对照全文 →