Qwen3-VL 技术报告

Qwen3-VL Technical Report

白金泽 Jinze Bai · · 2025-11-26 · arXiv:2511.21631 ↗ · 被引 1491

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

我们推出了 Qwen3-VL,这是 Qwen 系列迄今为止最强大的视觉语言模型,在广泛的多模态基准测试中均展现出卓越性能。它原生支持高达 256K 令牌的交错上下文,无缝整合文本、图像和视频。模型系列包括密集型(2B/4B/8B/32B)和混合专家(30B-A3B/235B-A22B)两种变体,以适应不同的延迟-质量权衡。Qwen3-VL 提供三大核心支柱:(i)显著增强的纯文本理解能力,在多个情况下超越同类纯文本骨干模型;(ii)稳健的长上下文理解能力,具备原生 256K 令牌窗口,适用于文本和交错的多种模态输入,能够忠实保持、检索和交叉引用长文档和视频中的信息;(iii)先进的单图像、多图像和视频任务多模态推理能力,在 MMMU 和视觉数学基准(如 MathVista 和 MathVision)等综合评估中表现领先。在架构方面,我们引入了三项关键升级:(i)增强的交错 MRoPE,用于更强的跨图像和视频的时空建模;(ii)深度融合 DeepStack,有效利用多级 ViT 特征来加强视觉-语言对齐;(iii)基于文本的视频时间对齐,从 T-RoPE 演进为显式文本时间戳对齐,实现更精确的时间定位。在可比的令牌预算和延迟约束下,Qwen3-VL 在密集架构和混合专家(MoE)架构中均实现了优越性能。我们期待 Qwen3-VL 能够作为图像基础推理、智能体决策以及多模态代码智能的基础引擎,在实际工作流程中发挥作用。

We introduce Qwen3-VL, the most capable vision-language model in the Qwen series to date, achieving superior performance across a broad range of multimodal benchmarks. It natively supports interleaved contexts of up to 256K tokens, seamlessly integrating text, images, and video. The model family includes both dense (2B/4B/8B/32B) and mixture-of-experts (30B-A3B/235B-A22B) variants to accommodate diverse latency-quality trade-offs. Qwen3-VL delivers three core pillars: (i) markedly stronger pure-text understanding, surpassing comparable text-only backbones in several cases; (ii) robust long-context comprehension with a native 256K-token window for both text and interleaved multimodal inputs, enabling faithful retention, retrieval, and cross-referencing across long documents and videos; and (iii) advanced multimodal reasoning across single-image, multi-image, and video tasks, demonstrating leading performance on comprehensive evaluations such as MMMU and visual-math benchmarks (e.g., MathVista and MathVision). Architecturally, we introduce three key upgrades: (i) an enhanced interleaved-MRoPE for stronger spatial-temporal modeling across images and video; (ii) DeepStack integration, which effectively leverages multi-level ViT features to tighten vision-language alignment; and (iii) text-based time alignment for video, evolving from T-RoPE to explicit textual timestamp alignment for more precise temporal grounding. Under comparable token budgets and latency constraints, Qwen3-VL achieves superior performance in both dense and Mixture-of-Experts (MoE) architectures. We envision Qwen3-VL serving as a foundational engine for image-grounded reasoning, agentic decision-making, and multimodal code intelligence in real-world workflows.

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 3)

阅读逐段中英对照全文 →