Janus-Pro:通过数据和模型扩展实现统一的多模态理解与生成

Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling

深度求索 DeepSeek-AI · DeepSeek · 2025-01-29 · arXiv:2501.17811 ↗ · 被引 775

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

在这项工作中,我们介绍了 Janus-Pro,这是先前工作 Janus 的进阶版本。具体来说,Janus-Pro 包含了(1)优化的训练策略,(2)扩展的训练数据,以及(3)更大规模的模型。通过这些改进,Janus-Pro 在多模态理解和文本到图像指令跟随能力方面取得了显著进步,同时增强了文本到图像生成的稳定性。我们希望这项工作能激发该领域的进一步探索。代码和模型已公开。

In this work, we introduce Janus-Pro, an advanced version of the previous work Janus. Specifically, Janus-Pro incorporates (1) an optimized training strategy, (2) expanded training data, and (3) scaling to larger model size. With these improvements, Janus-Pro achieves significant advancements in both multimodal understanding and text-to-image instruction-following capabilities, while also enhancing the stability of text-to-image generation. We hope this work will inspire further exploration in the field. Code and models are publicly available.

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 11)

阅读逐段中英对照全文 →