Llama 3 模型群

The Llama 3 Herd of Models

Aaron Grattafiori Aaron Grattafiori · · 2024-07-31 · arXiv:2407.21783 ↗ · 被引 17476

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

现代人工智能系统由基础模型驱动。本文介绍了一组新的基础模型,称为 Llama 3。这是一个语言模型群,原生支持多语言、编码、推理和工具使用。我们的最大模型是一个稠密 Transformer,拥有 405B 参数和高达 128K token 的上下文窗口。本文对 Llama 3 进行了广泛的实证评估。我们发现,Llama 3 在众多任务上提供了与 GPT-4 等领先语言模型相当的质量。我们公开发布了 Llama 3,包括 405B 参数语言模型的预训练和后训练版本,以及用于输入和输出安全的 Llama Guard 3 模型。本文还介绍了通过组合方法将图像、视频和语音能力集成到 Llama 3 中的实验结果。我们观察到,这种方法在图像、视频和语音识别任务上表现出与最先进技术相当的水平。由此产生的模型尚未广泛发布,因为它们仍在开发中。

Modern artificial intelligence (AI) systems are powered by foundation models. This paper presents a new set of foundation models, called Llama 3. It is a herd of language models that natively support multilinguality, coding, reasoning, and tool usage. Our largest model is a dense Transformer with 405B parameters and a context window of up to 128K tokens. This paper presents an extensive empirical evaluation of Llama 3. We find that Llama 3 delivers comparable quality to leading language models such as GPT-4 on a plethora of tasks. We publicly release Llama 3, including pre-trained and post-trained versions of the 405B parameter language model and our Llama Guard 3 model for input and output safety. The paper also presents the results of experiments in which we integrate image, video, and speech capabilities into Llama 3 via a compositional approach. We observe this approach performs competitively with the state-of-the-art on image, video, and speech recognition tasks. The resulting models are not yet being broadly released as they are still under development.

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 14)

阅读逐段中英对照全文 →