Qwen3.5-Omni 技术报告

Qwen3.5-Omni Technical Report

林俊旸 Junyang Lin · Alibaba Qwen · 2026-04-17 · arXiv:2604.15804 ↗ · 被引 71

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

本文介绍了 Qwen3.5-Omni,这是 Qwen-Omni 模型系列的最新进展。相比前代,Qwen3.5-Omni 规模扩展至数千亿参数,支持 256k 上下文长度。通过利用包含异构文本-视觉对和超过 1 亿小时音视频内容的大规模数据集,该模型展现出强大的全模态能力。Qwen3.5-Omni-plus 在 215 个音频及音视频理解、推理和交互子任务与基准测试中达到 SOTA,在关键音频任务上超越 Gemini-3.1 Pro,并在综合音视频理解上与之持平。架构上,Qwen3.5-Omni 在 Thinker 和 Talker 中均采用混合注意力混合专家(MoE)框架,实现高效长序列推理。该模型支持复杂交互,可理解超过 10 小时的音频和 400 秒的 720P 视频(1 FPS)。为解决流式语音合成中因文本与语音分词器编码效率差异导致的固有不稳定和不自然问题,我们引入了 ARIA。ARIA 动态对齐文本和语音单元,显著提升对话语音的稳定性和韵律,且延迟影响极小。此外,Qwen3.5-Omni 拓展了语言边界,支持 10 种语言的多语言理解和语音生成,并带有类人情感细微差别。最后,Qwen3.5-Omni 展现出卓越的音视频定位能力,可生成具有精确时间同步和自动场景分割的脚本级结构化字幕。值得注意的是,我们观察到全模态模型中出现了一种新能力:直接基于音视频指令进行编码,我们称之为音视频氛围编码。

In this work, we present Qwen3.5-Omni, the latest advancement in the Qwen-Omni model family. Representing a significant evolution over its predecessor, Qwen3.5-Omni scales to hundreds of billions of parameters and supports a 256k context length. By leveraging a massive dataset comprising heterogeneous text-vision pairs and over 100 million hours of audio-visual content, the model demonstrates robust omni-modality capabilities. Qwen3.5-Omni-plus achieves SOTA results across 215 audio and audio-visual understanding, reasoning, and interaction subtasks and benchmarks, surpassing Gemini-3.1 Pro in key audio tasks and matching it in comprehensive audio-visual understanding. Architecturally, Qwen3.5-Omni employs a Hybrid Attention Mixture-of-Experts (MoE) framework for both Thinker and Talker, enabling efficient long-sequence inference. The model facilitates sophisticated interaction, supporting over 10 hours of audio understanding and 400 seconds of 720P video (at 1 FPS). To address the inherent instability and unnaturalness in streaming speech synthesis, often caused by encoding efficiency discrepancies between text and speech tokenizers, we introduce ARIA. ARIA dynamically aligns text and speech units, significantly enhancing the stability and prosody of conversational speech with minimal latency impact. Furthermore, Qwen3.5-Omni expands linguistic boundaries, supporting multilingual understanding and speech generation across 10 languages with human-like emotional nuance. Finally, Qwen3.5-Omni exhibits superior audio-visual grounding capabilities, generating script-level structured captions with precise temporal synchronization and automated scene segmentation. Remarkably, we observed the emergence of a new capability in omnimodal models: directly performing coding based on audio-visual instructions, which we call Audio-Visual Vibe Coding.

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 17)

阅读逐段中英对照全文 →