Step-Audio 2 技术报告

Step-Audio 2 Technical Report

阶跃星辰 StepFun · StepFun · 2025-07-22 · arXiv:2507.16632 ↗ · 被引 109

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

本文介绍了 Step-Audio 2,一个为工业级音频理解和语音对话设计的端到端多模态大语言模型。通过集成潜在音频编码器和以推理为中心的强化学习(RL),Step-Audio 2 在自动语音识别(ASR)和音频理解方面取得了令人瞩目的性能。为了实现真正的端到端语音对话,Step-Audio 2 将离散音频令牌的生成融入语言建模,显著增强了对副语言信息(如说话风格和情感)的响应能力。为了有效利用现实世界数据中的丰富文本和声学知识,Step-Audio 2 集成了检索增强生成(RAG),并能够调用外部工具(如网络搜索)以减少幻觉,以及音频搜索以切换音色。经过数百万小时的语音和音频数据训练,Step-Audio 2 在各种对话场景中展现出智能和表现力。评估结果表明,与其他开源和商业解决方案相比,Step-Audio 2 在多种音频理解和对话基准测试中达到了最先进的性能。更多信息请访问 https://github.com/stepfun-ai/Step-Audio2。

This paper presents Step-Audio 2, an end-to-end multi-modal large language model designed for industry-strength audio understanding and speech conversation. By integrating a latent audio encoder and reasoning-centric reinforcement learning (RL), Step-Audio 2 achieves promising performance in automatic speech recognition (ASR) and audio understanding. To facilitate genuine end-to-end speech conversation, Step-Audio 2 incorporates the generation of discrete audio tokens into language modeling, significantly enhancing its responsiveness to paralinguistic information such as speaking styles and emotions. To effectively leverage the rich textual and acoustic knowledge in real-world data, Step-Audio 2 integrates retrieval-augmented generation (RAG) and is able to call external tools such as web search to mitigate hallucination and audio search to switch timbres. Trained on millions of hours of speech and audio data, Step-Audio 2 delivers intelligence and expressiveness across diverse conversational scenarios. Evaluation results demonstrate that Step-Audio 2 achieves state-of-the-art performance on various audio understanding and conversational benchmarks compared to other open-source and commercial solutions. Please visit https://github.com/stepfun-ai/Step-Audio2 for more information.

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 17)

阅读逐段中英对照全文 →