Qwen2-Audio Technical Report
打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→我们介绍了 Qwen-Audio 的最新进展,即名为 Qwen2-Audio 的大规模音频语言模型,它能够接受各种音频信号输入,并根据语音指令进行音频分析或直接生成文本回复。与复杂的层次化标签不同,我们通过使用自然语言提示来简化不同数据和任务的预训练过程,并进一步扩展了数据量。我们增强了 Qwen2-Audio 的指令跟随能力,并实现了两种不同的音频交互模式:语音聊天和音频分析。在语音聊天模式下,用户可以无需文本输入,自由地与 Qwen2-Audio 进行语音交互。在音频分析模式下,用户可以在交互过程中提供音频和文本指令进行分析。注意,我们未使用任何系统提示来切换语音聊天和音频分析模式。Qwen2-Audio 能够智能地理解音频内容,并遵循语音命令做出适当响应。例如,在同时包含声音、多说话人对话和语音命令的音频片段中,Qwen2-Audio 可以直接理解命令并提供对音频的解释和响应。此外,DPO 优化了模型在事实性和遵循期望行为方面的性能。根据 AIR-Bench 的评估结果,Qwen2-Audio 在专注于音频中心指令跟随能力的测试中优于之前的 SOTA,如 Gemini-1.5-pro。Qwen2-Audio 已开源,旨在推动多模态语言社区的发展。
We introduce the latest progress of Qwen-Audio, a large-scale audio-language model called Qwen2-Audio, which is capable of accepting various audio signal inputs and performing audio analysis or direct textual responses with regard to speech instructions. In contrast to complex hierarchical tags, we have simplified the pre-training process by utilizing natural language prompts for different data and tasks, and have further expanded the data volume. We have boosted the instruction-following capability of Qwen2-Audio and implemented two distinct audio interaction modes for voice chat and audio analysis. In the voice chat mode, users can freely engage in voice interactions with Qwen2-Audio without text input. In the audio analysis mode, users could provide audio and text instructions for analysis during the interaction. Note that we do not use any system prompts to switch between voice chat and audio analysis modes. Qwen2-Audio is capable of intelligently comprehending the content within audio and following voice commands to respond appropriately. For instance, in an audio segment that simultaneously contains sounds, multi-speaker conversations, and a voice command, Qwen2-Audio can directly understand the command and provide an interpretation and response to the audio. Additionally, DPO has optimized the model's performance in terms of factuality and adherence to desired behavior. According to the evaluation results from AIR-Bench, Qwen2-Audio outperformed previous SOTAs, such as Gemini-1.5-pro, in tests focused on audio-centric instruction-following capabilities. Qwen2-Audio is open-sourced with the aim of fostering the advancement of the multi-modal language community.