Qwen2.5-VL Technical Report
打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→我们推出了 Qwen2.5-VL,这是 Qwen 视觉语言系列的最新旗舰模型,在基础能力和创新功能方面均取得了显著进展。Qwen2.5-VL 通过增强的视觉识别、精确的物体定位、稳健的文档解析和长视频理解,在理解和交互世界方面实现了重大飞跃。其一个突出特点是能够使用边界框或点精确地定位物体。它能够从发票、表单和表格中提取稳健的结构化数据,并对图表、示意图和布局进行详细分析。为了处理复杂输入,Qwen2.5-VL 引入了动态分辨率处理和绝对时间编码,使其能够处理不同尺寸的图像和长达数小时的视频,并实现秒级事件定位。这使得模型能够原生感知空间尺度和时间动态,而无需依赖传统的归一化技术。通过从头训练原生动态分辨率的视觉 Transformer(ViT)并结合窗口注意力机制,我们在保持原生分辨率的同时降低了计算开销。因此,Qwen2.5-VL 不仅在静态图像和文档理解方面表现出色,还能作为交互式视觉代理,在真实场景中执行推理、工具使用和任务执行,例如操作计算机和移动设备。Qwen2.5-VL 提供三种尺寸,满足从边缘 AI 到高性能计算的不同应用场景。旗舰版 Qwen2.5-VL-72B 模型与 GPT-4o 和 Claude 3.5 Sonnet 等最先进模型相媲美,尤其在文档和图表理解方面表现卓越。此外,Qwen2.5-VL 保持了强大的语言性能,保留了 Qwen2.5 大语言模型的核心语言能力。
We introduce Qwen2.5-VL, the latest flagship model of Qwen vision-language series, which demonstrates significant advancements in both foundational capabilities and innovative functionalities. Qwen2.5-VL achieves a major leap forward in understanding and interacting with the world through enhanced visual recognition, precise object localization, robust document parsing, and long-video comprehension. A standout feature of Qwen2.5-VL is its ability to localize objects using bounding boxes or points accurately. It provides robust structured data extraction from invoices, forms, and tables, as well as detailed analysis of charts, diagrams, and layouts. To handle complex inputs, Qwen2.5-VL introduces dynamic resolution processing and absolute time encoding, enabling it to process images of varying sizes and videos of extended durations (up to hours) with second-level event localization. This allows the model to natively perceive spatial scales and temporal dynamics without relying on traditional normalization techniques. By training a native dynamic-resolution Vision Transformer (ViT) from scratch and incorporating Window Attention, we reduce computational overhead while maintaining native resolution. As a result, Qwen2.5-VL excels not only in static image and document understanding but also as an interactive visual agent capable of reasoning, tool usage, and task execution in real-world scenarios such as operating computers and mobile devices. Qwen2.5-VL is available in three sizes, addressing diverse use cases from edge AI to high-performance computing. The flagship Qwen2.5-VL-72B model matches state-of-the-art models like GPT-4o and Claude 3.5 Sonnet, particularly excelling in document and diagram understanding. Additionally, Qwen2.5-VL maintains robust linguistic performance, preserving the core language competencies of the Qwen2.5 LLM.