Pixtral 12B:120 亿参数多模态语言模型

Pixtral 12B

蒋启天 Albert Q. Jiang · Mistral AI · 2024-10-09 · arXiv:2410.07073 ↗ · 被引 157

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

我们推出了 Pixtral-12B,一个 120 亿参数的多模态语言模型。Pixtral-12B 经过训练,能够理解自然图像和文档,在各种多模态基准测试中取得了领先性能,超越了许多更大的模型。与许多开源模型不同,Pixtral 在其规模上也是一款尖端的文本模型,并且不会为了在多模态任务中表现出色而牺牲自然语言性能。Pixtral 使用了一个从头开始训练的新视觉编码器,使其能够以自然分辨率和宽高比处理图像。这为用户在处理图像时使用的令牌数量提供了灵活性。Pixtral 还能够在 128K 令牌的长上下文窗口中处理任意数量的图像。Pixtral 12B 显著优于其他类似规模的开源模型(Llama-3.2 11B 和 Qwen-2-VL 7B)。它还优于更大的开源模型,如 Llama-3.2 90B,同时体积小 7 倍。我们进一步贡献了一个开源基准测试 MM-MT-Bench,用于评估实际场景中的视觉语言模型,并为多模态 LLM 的标准化评估协议提供了详细分析和代码。Pixtral-12B 在 Apache 2.0 许可下发布。

We introduce Pixtral-12B, a 12--billion-parameter multimodal language model. Pixtral-12B is trained to understand both natural images and documents, achieving leading performance on various multimodal benchmarks, surpassing a number of larger models. Unlike many open-source models, Pixtral is also a cutting-edge text model for its size, and does not compromise on natural language performance to excel in multimodal tasks. Pixtral uses a new vision encoder trained from scratch, which allows it to ingest images at their natural resolution and aspect ratio. This gives users flexibility on the number of tokens used to process an image. Pixtral is also able to process any number of images in its long context window of 128K tokens. Pixtral 12B substanially outperforms other open models of similar sizes (Llama-3.2 11B \& Qwen-2-VL 7B). It also outperforms much larger open models like Llama-3.2 90B while being 7x smaller. We further contribute an open-source benchmark, MM-MT-Bench, for evaluating vision-language models in practical scenarios, and provide detailed analysis and code for standardized evaluation protocols for multimodal LLMs. Pixtral-12B is released under Apache 2.0 license.

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 8)

阅读逐段中英对照全文 →