Nemotron 3 Super:面向智能体推理的开放、高效混合专家 Mamba-Transformer 模型

Nemotron 3 Super: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning

Akhiad Bercovich Akhiad Bercovich · NVIDIA · 2026-04-14 · arXiv:2604.12374 ↗ · 被引 16

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

我们描述了 Nemotron 3 Super 的预训练、后训练和量化过程,这是一个 1200 亿(活跃 120 亿)参数的混合 Mamba-注意力混合专家模型。Nemotron 3 Super 是 Nemotron 3 系列中首个:1)采用 NVFP4 预训练,2)利用 LatentMoE(一种新的混合专家架构,优化了每 FLOP 精度和每参数精度),3)包含 MTP 层,通过原生推测解码加速推理。我们在 25 万亿 token 上预训练了 Nemotron 3 Super,随后使用监督微调(SFT)和强化学习(RL)进行后训练。最终模型支持高达 100 万上下文长度,在常见基准测试上达到相当精度,同时推理吞吐量分别比 GPT-OSS-120B 和 Qwen3.5-122B 高出 2.2 倍和 7.5 倍。Nemotron 3 Super 的数据集以及基础、后训练和量化检查点已在 HuggingFace 上开源。

We describe the pre-training, post-training, and quantization of Nemotron 3 Super, a 120 billion (active 12 billion) parameter hybrid Mamba-Attention Mixture-of-Experts model. Nemotron 3 Super is the first model in the Nemotron 3 family to 1) be pre-trained in NVFP4, 2) leverage LatentMoE, a new Mixture-of-Experts architecture that optimizes for both accuracy per FLOP and accuracy per parameter, and 3) include MTP layers for inference acceleration through native speculative decoding. We pre-trained Nemotron 3 Super on 25 trillion tokens followed by post-training using supervised fine tuning (SFT) and reinforcement learning (RL). The final model supports up to 1M context length and achieves comparable accuracy on common benchmarks, while also achieving up to 2.2x and 7.5x higher inference throughput compared to GPT-OSS-120B and Qwen3.5-122B, respectively. Nemotron 3 Super datasets, along with the base, post-trained, and quantized checkpoints, are open-sourced on HuggingFace.

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 23)

阅读逐段中英对照全文 →