Aria:一个开放的多模态原生混合专家模型

Aria: An Open Multimodal Native Mixture-of-Experts Model

李俊男 Junnan Li · Rhymes AI · 2024-10-08 · arXiv:2410.05993 ↗ · 被引 139

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

信息以多种模态呈现。多模态原生 AI 模型对于整合现实世界信息并提供全面理解至关重要。尽管存在专有的多模态原生模型,但它们缺乏开放性,这给采用带来了障碍,更不用说适应了。为了填补这一空白,我们引入了 Aria,一个开放的多模态原生模型,在广泛的多模态、语言和编码任务中具有最佳性能。Aria 是一个混合专家模型,每个视觉令牌和文本令牌分别有 3.9B 和 3.5B 激活参数。它在各种多模态任务上优于 Pixtral-12B 和 Llama3.2-11B,并与最佳专有模型竞争。我们按照 4 阶段流程从头开始预训练 Aria,逐步使模型具备语言理解、多模态理解、长上下文窗口和指令遵循的强大能力。我们开源了模型权重以及一个代码库,以促进 Aria 在实际应用中的轻松采用和适应。

Information comes in diverse modalities. Multimodal native AI models are essential to integrate real-world information and deliver comprehensive understanding. While proprietary multimodal native models exist, their lack of openness imposes obstacles for adoptions, let alone adaptations. To fill this gap, we introduce Aria, an open multimodal native model with best-in-class performance across a wide range of multimodal, language, and coding tasks. Aria is a mixture-of-expert model with 3.9B and 3.5B activated parameters per visual token and text token, respectively. It outperforms Pixtral-12B and Llama3.2-11B, and is competitive against the best proprietary models on various multimodal tasks. We pre-train Aria from scratch following a 4-stage pipeline, which progressively equips the model with strong capabilities in language understanding, multimodal understanding, long context window, and instruction following. We open-source the model weights along with a codebase that facilitates easy adoptions and adaptations of Aria in real-world applications.

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 13)

阅读逐段中英对照全文 →