OLMo:加速语言模型科学

OLMo: Accelerating the Science of Language Models

Dirk Groeneveld Dirk Groeneveld · Allen Institute for AI · 2024-02-01 · arXiv:2402.00838 ↗ · 被引 683

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

语言模型在自然语言处理研究和商业产品中已无处不在。随着其商业重要性激增,最强大的模型变得封闭,隐藏在专有接口之后,其训练数据、架构和开发的重要细节未公开。鉴于这些细节在科学研究这些模型(包括其偏见和潜在风险)中的重要性,我们认为研究社区必须能够访问强大、真正开放的语言模型。为此,我们构建了 OLMo,一个具有竞争力的、真正开放的语言模型,以实现对语言模型的科学研究。与大多数仅发布模型权重和推理代码的先前努力不同,我们同时发布了 OLMo 的开放训练数据以及训练和评估代码。我们希望这一发布将赋能开放研究社区,并激发新一轮创新。

Language models (LMs) have become ubiquitous in both NLP research and in commercial product offerings. As their commercial importance has surged, the most powerful models have become closed off, gated behind proprietary interfaces, with important details of their training data, architectures, and development undisclosed. Given the importance of these details in scientifically studying these models, including their biases and potential risks, we believe it is essential for the research community to have access to powerful, truly open LMs. To this end, we have built OLMo, a competitive, truly Open Language Model, to enable the scientific study of language models. Unlike most prior efforts that have only released model weights and inference code, we release OLMo alongside open training data and training and evaluation code. We hope this release will empower the open research community and inspire a new wave of innovation.

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 22)

阅读逐段中英对照全文 →