CosyVoice 3:通过规模扩展和后训练实现野外语音生成

CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training

杜志浩 Zhihao Du · Alibaba · 2025-05-23 · arXiv:2505.17589 ↗ · 被引 164

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

在之前的工作中,我们介绍了一个可扩展的流式语音合成模型 CosyVoice 2,它集成了大语言模型(LLM)和块感知流匹配(FM)模型,实现了低延迟的双流式语音合成和与人类相当的质量。尽管取得了这些进展,CosyVoice 2 在语言覆盖、领域多样性、数据量、文本格式和后训练技术方面仍存在局限性。在本文中,我们提出了 CosyVoice 3,这是一个改进的模型,专为野外零样本多语言语音合成而设计,在内容一致性、说话人相似性和韵律自然度方面超越了其前身。CosyVoice 3 的关键特性包括:1)一种新颖的语音分词器,通过监督多任务训练(包括自动语音识别、语音情感识别、语言识别、音频事件检测和说话人分析)来提高韵律自然度。2)一种新的可微分奖励模型,用于后训练,不仅适用于 CosyVoice 3,也适用于其他基于 LLM 的语音合成模型。3)数据集规模扩展:训练数据从一万小时扩展到一百万小时,涵盖 9 种语言和 18 种中文方言,涉及各种领域和文本格式。4)模型规模扩展:模型参数从 5 亿增加到 15 亿,由于更大的模型容量,在我们的多语言基准测试中性能得到提升。这些进展显著推动了野外语音合成的发展。我们鼓励读者在 https://funaudiollm.github.io/cosyvoice3 上收听演示。

In our prior works, we introduced a scalable streaming speech synthesis model, CosyVoice 2, which integrates a large language model (LLM) and a chunk-aware flow matching (FM) model, and achieves low-latency bi-streaming speech synthesis and human-parity quality. Despite these advancements, CosyVoice 2 exhibits limitations in language coverage, domain diversity, data volume, text formats, and post-training techniques. In this paper, we present CosyVoice 3, an improved model designed for zero-shot multilingual speech synthesis in the wild, surpassing its predecessor in content consistency, speaker similarity, and prosody naturalness. Key features of CosyVoice 3 include: 1) A novel speech tokenizer to improve prosody naturalness, developed via supervised multi-task training, including automatic speech recognition, speech emotion recognition, language identification, audio event detection, and speaker analysis. 2) A new differentiable reward model for post-training applicable not only to CosyVoice 3 but also to other LLM-based speech synthesis models. 3) Dataset Size Scaling: Training data is expanded from ten thousand hours to one million hours, encompassing 9 languages and 18 Chinese dialects across various domains and text formats. 4) Model Size Scaling: Model parameters are increased from 0.5 billion to 1.5 billion, resulting in enhanced performance on our multilingual benchmark due to the larger model capacity. These advancements contribute significantly to the progress of speech synthesis in the wild. We encourage readers to listen to the demo at https://funaudiollm.github.io/cosyvoice3.

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 24)

阅读逐段中英对照全文 →