Step-Audio:智能语音交互中的统一理解与生成

Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction

阶跃星辰 StepFun · StepFun · 2025-02-17 · arXiv:2502.11946 ↗ · 被引 115

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

实时语音交互作为人机协作的基本接口,潜力巨大。然而,当前开源模型面临语音数据采集成本高、动态控制能力弱、智能水平有限等挑战。为此,本文提出 Step-Audio,首个生产级开源解决方案。主要贡献包括:1)130B 参数统一语音-文本多模态模型,实现统一理解与生成,并开源 Step-Audio-Chat 版本;2)生成式语音数据引擎,建立经济实惠的语音克隆框架,并通过蒸馏开源轻量级 Step-Audio-TTS-3B 模型;3)指令驱动的精细控制系统,支持方言、情感、歌唱和 RAP 的动态调整;4)增强的认知架构,集成工具调用和角色扮演能力,有效管理复杂任务。基于新提出的 StepEval-Audio-360 评估基准,Step-Audio 在人工评估中达到最先进性能,尤其在指令遵循方面。在 LLaMA Question 等开源基准上,平均性能提升 9.3%,展示了我们对推进开源多模态语言技术发展的承诺。代码和模型见 https://github.com/stepfun-ai/Step-Audio。

Real-time speech interaction, serving as a fundamental interface for human-machine collaboration, holds immense potential. However, current open-source models face limitations such as high costs in voice data collection, weakness in dynamic control, and limited intelligence. To address these challenges, this paper introduces Step-Audio, the first production-ready open-source solution. Key contributions include: 1) a 130B-parameter unified speech-text multi-modal model that achieves unified understanding and generation, with the Step-Audio-Chat version open-sourced; 2) a generative speech data engine that establishes an affordable voice cloning framework and produces the open-sourced lightweight Step-Audio-TTS-3B model through distillation; 3) an instruction-driven fine control system enabling dynamic adjustments across dialects, emotions, singing, and RAP; 4) an enhanced cognitive architecture augmented with tool calling and role-playing abilities to manage complex tasks effectively. Based on our new StepEval-Audio-360 evaluation benchmark, Step-Audio achieves state-of-the-art performance in human evaluations, especially in terms of instruction following. On open-source benchmarks like LLaMA Question, shows 9.3% average performance improvement, demonstrating our commitment to advancing the development of open-source multi-modal language technologies. Our code and models are available at https://github.com/stepfun-ai/Step-Audio.

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 20)

阅读逐段中英对照全文 →