ERNIE 5.0 技术报告

ERNIE 5.0 Technical Report

王海峰 Haifeng Wang · Baidu · 2026-02-04 · arXiv:2602.04705 ↗ · 被引 8

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

本报告介绍了 ERNIE 5.0,一个原生自回归基础模型,专为跨文本、图像、视频和音频的统一多模态理解与生成而设计。所有模态在统一的下一组令牌预测目标下从头开始训练,基于具有模态无关专家路由的超稀疏混合专家(MoE)架构。为了应对不同资源约束下大规模部署的实际挑战,ERNIE 5.0 采用了一种新颖的弹性训练范式。在单次预训练运行中,模型学习了一系列具有不同深度、专家容量和路由稀疏性的子模型,从而在内存或时间受限的场景中实现性能、模型大小和推理延迟之间的灵活权衡。此外,我们系统地解决了将强化学习扩展到统一基础模型的挑战,从而保证了在超稀疏 MoE 架构和多样化多模态设置下的高效稳定后训练。大量实验表明,ERNIE 5.0 在多种模态上实现了强大且均衡的性能。据我们所知,在公开披露的模型中,ERNIE 5.0 代表了首个支持多模态理解与生成的生产级万亿参数统一自回归模型。为了促进进一步研究,我们提供了统一模型中模态无关专家路由的详细可视化,以及弹性训练的全面实证分析,旨在为社区提供深刻的见解。

In this report, we introduce ERNIE 5.0, a natively autoregressive foundation model desinged for unified multimodal understanding and generation across text, image, video, and audio. All modalities are trained from scratch under a unified next-group-of-tokens prediction objective, based on an ultra-sparse mixture-of-experts (MoE) architecture with modality-agnostic expert routing. To address practical challenges in large-scale deployment under diverse resource constraints, ERNIE 5.0 adopts a novel elastic training paradigm. Within a single pre-training run, the model learns a family of sub-models with varying depths, expert capacities, and routing sparsity, enabling flexible trade-offs among performance, model size, and inference latency in memory- or time-constrained scenarios. Moreover, we systematically address the challenges of scaling reinforcement learning to unified foundation models, thereby guaranteeing efficient and stable post-training under ultra-sparse MoE architectures and diverse multimodal settings. Extensive experiments demonstrate that ERNIE 5.0 achieves strong and balanced performance across multiple modalities. To the best of our knowledge, among publicly disclosed models, ERNIE 5.0 represents the first production-scale realization of a trillion-parameter unified autoregressive model that supports both multimodal understanding and generation. To facilitate further research, we present detailed visualizations of modality-agnostic expert routing in the unified model, alongside comprehensive empirical analysis of elastic training, aiming to offer profound insights to the community.

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 25)

阅读逐段中英对照全文 →