DeepSeek-V3.2:推动开放大型语言模型的前沿

DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models

深度求索 DeepSeek-AI · DeepSeek · 2025-12-02 · arXiv:2512.02556 ↗ · 被引 599

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

我们推出了 DeepSeek-V3.2,这是一个兼具高计算效率与卓越推理及智能体性能的模型。其关键技术突破如下:(1)DeepSeek 稀疏注意力(DSA):我们引入了 DSA,一种高效的注意力机制,在长上下文场景中大幅降低计算复杂度,同时保持模型性能。(2)可扩展强化学习框架:通过实施稳健的强化学习协议并扩展后训练计算量,DeepSeek-V3.2 的性能与 GPT-5 相当。值得注意的是,我们的高计算变体 DeepSeek-V3.2-Speciale 超越了 GPT-5,推理能力与 Gemini-3.0-Pro 持平,在 2025 年国际数学奥林匹克竞赛(IMO)和国际信息学奥林匹克竞赛(IOI)中均获得金牌成绩。(3)大规模智能体任务合成流水线:为了将推理融入工具使用场景,我们开发了一种新颖的合成流水线,能够系统性地大规模生成训练数据。该方法支持可扩展的智能体后训练,在复杂交互环境中显著提升了泛化能力和指令遵循的鲁棒性。

We introduce DeepSeek-V3.2, a model that harmonizes high computational efficiency with superior reasoning and agent performance. The key technical breakthroughs of DeepSeek-V3.2 are as follows: (1) DeepSeek Sparse Attention (DSA): We introduce DSA, an efficient attention mechanism that substantially reduces computational complexity while preserving model performance in long-context scenarios. (2) Scalable Reinforcement Learning Framework: By implementing a robust reinforcement learning protocol and scaling post-training compute, DeepSeek-V3.2 performs comparably to GPT-5. Notably, our high-compute variant, DeepSeek-V3.2-Speciale, surpasses GPT-5 and exhibits reasoning proficiency on par with Gemini-3.0-Pro, achieving gold-medal performance in both the 2025 International Mathematical Olympiad (IMO) and the International Olympiad in Informatics (IOI). (3) Large-Scale Agentic Task Synthesis Pipeline: To integrate reasoning into tool-use scenarios, we developed a novel synthesis pipeline that systematically generates training data at scale. This methodology facilitates scalable agentic post-training, yielding substantial improvements in generalization and instruction-following robustness within complex, interactive environments.

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 17)

阅读逐段中英对照全文 →