Gemma 2:在实用规模上改进开放语言模型

Gemma 2: Improving Open Language Models at a Practical Size

Morgane Riviere Morgane Riviere · Google DeepMind · 2024-07-31 · arXiv:2408.00118 ↗ · 被引 2150

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

在这项工作中,我们推出了 Gemma 2,这是 Gemma 系列轻量级、最先进开放模型的新成员,参数规模从 20 亿到 270 亿不等。在这个新版本中,我们对 Transformer 架构应用了几项已知的技术改进,例如交错局部-全局注意力(Beltagy 等人,2020a)和分组查询注意力(Ainslie 等人,2023)。我们还使用知识蒸馏(Hinton 等人,2015)而非下一个词预测来训练 2B 和 9B 模型。由此产生的模型在其规模上提供了最佳性能,甚至为比它们大 2-3 倍的模型提供了有竞争力的替代方案。我们将所有模型发布给社区。

In this work, we introduce Gemma 2, a new addition to the Gemma family of lightweight, state-of-the-art open models, ranging in scale from 2 billion to 27 billion parameters. In this new version, we apply several known technical modifications to the Transformer architecture, such as interleaving local-global attentions (Beltagy et al., 2020a) and group-query attention (Ainslie et al., 2023). We also train the 2B and 9B models with knowledge distillation (Hinton et al., 2015) instead of next token prediction. The resulting models deliver the best performance for their size, and even offer competitive alternatives to models that are 2-3 times bigger. We release all our models to the community.

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 22)

阅读逐段中英对照全文 →