通过大规模弱监督实现鲁棒语音识别

Robust Speech Recognition via Large-Scale Weak Supervision

亚历克·拉德福德 Alec Radford · OpenAI · 2022-12-06 · arXiv:2212.04356 ↗ · 被引 7628

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

我们研究了仅通过预测互联网上大量音频转录本来训练的语音处理系统的能力。当扩展到 68 万小时的多语言和多任务监督时,得到的模型在标准基准测试中表现良好,通常与先前的全监督结果相当,但处于零样本迁移设置,无需任何微调。与人类相比,这些模型接近其准确性和鲁棒性。我们正在发布模型和推理代码,作为进一步研究鲁棒语音处理的基础。

We study the capabilities of speech processing systems trained simply to predict large amounts of transcripts of audio on the internet. When scaled to 680,000 hours of multilingual and multitask supervision, the resulting models generalize well to standard benchmarks and are often competitive with prior fully supervised results but in a zero-shot transfer setting without the need for any fine-tuning. When compared to humans, the models approach their accuracy and robustness. We are releasing models and inference code to serve as a foundation for further work on robust speech processing.

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 23)

阅读逐段中英对照全文 →