LLaMA:开放且高效的基础语言模型

LLaMA: Open and Efficient Foundation Language Models

Aaron Grattafiori Aaron Grattafiori · Meta AI · 2023-02-27 · arXiv:2302.13971 ↗ · 被引 20615

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

我们介绍了 LLaMA,一系列参数规模从 7B 到 65B 的基础语言模型。我们在数万亿个 token 上训练模型,并证明仅使用公开可用的数据集,而不依赖专有和不可访问的数据集,就能训练出最先进的模型。特别是,LLaMA-13B 在大多数基准测试上优于 GPT-3(175B),而 LLaMA-65B 与最佳模型 Chinchilla-70B 和 PaLM-540B 不相上下。我们将所有模型发布给研究社区。

We introduce LLaMA, a collection of foundation language models ranging from 7B to 65B parameters. We train our models on trillions of tokens, and show that it is possible to train state-of-the-art models using publicly available datasets exclusively, without resorting to proprietary and inaccessible datasets. In particular, LLaMA-13B outperforms GPT-3 (175B) on most benchmarks, and LLaMA-65B is competitive with the best models, Chinchilla-70B and PaLM-540B. We release all our models to the research community.

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 25)

阅读逐段中英对照全文 →