DeepSeek LLM:以长期主义扩展开源语言模型

DeepSeek LLM: Scaling Open-Source Language Models with Longtermism

深度求索 DeepSeek-AI · DeepSeek · 2024-01-05 · arXiv:2401.02954 ↗ · 被引 785

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

开源大型语言模型(LLM)的快速发展令人瞩目。然而,先前文献中描述的缩放定律呈现出不同的结论,给 LLM 的扩展蒙上了阴影。我们深入研究了缩放定律,并提出了独特的发现,这些发现有助于在两种常用的开源配置(7B 和 67B)中扩展大规模模型。在缩放定律的指导下,我们推出了 DeepSeek LLM,这是一个致力于以长期视角推进开源语言模型的项目。为了支持预训练阶段,我们开发了一个数据集,目前包含 2 万亿个 token,并且还在持续扩展。我们进一步在 DeepSeek LLM 基础模型上进行了监督微调(SFT)和直接偏好优化(DPO),从而创建了 DeepSeek Chat 模型。我们的评估结果表明,DeepSeek LLM 67B 在各种基准测试上超越了 LLaMA-2 70B,特别是在代码、数学和推理领域。此外,开放式评估显示,DeepSeek LLM 67B Chat 的性能优于 GPT-3.5。

The rapid development of open-source large language models (LLMs) has been truly remarkable. However, the scaling law described in previous literature presents varying conclusions, which casts a dark cloud over scaling LLMs. We delve into the study of scaling laws and present our distinctive findings that facilitate scaling of large scale models in two commonly used open-source configurations, 7B and 67B. Guided by the scaling laws, we introduce DeepSeek LLM, a project dedicated to advancing open-source language models with a long-term perspective. To support the pre-training phase, we have developed a dataset that currently consists of 2 trillion tokens and is continuously expanding. We further conduct supervised fine-tuning (SFT) and Direct Preference Optimization (DPO) on DeepSeek LLM Base models, resulting in the creation of DeepSeek Chat models. Our evaluation results demonstrate that DeepSeek LLM 67B surpasses LLaMA-2 70B on various benchmarks, particularly in the domains of code, mathematics, and reasoning. Furthermore, open-ended evaluations reveal that DeepSeek LLM 67B Chat exhibits superior performance compared to GPT-3.5.

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 17)

阅读逐段中英对照全文 →