Aya 模型:一种指令微调的开源多语言语言模型

Aya Model: An Instruction Finetuned Open-Access Multilingual Language Model

Sara Hooker Sara Hooker · · 2024-02-12 · arXiv:2402.07827 ↗ · 被引 389

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

近期大语言模型(LLMs)的突破主要集中在少数数据丰富的语言上。如何将突破扩展到这些一流语言之外?我们的工作引入了 Aya,一个大规模多语言生成语言模型,能够执行 101 种语言的指令,其中超过 50%被认为是低资源语言。Aya 在大多数任务上优于 mT0 和 BLOOMZ,同时覆盖的语言数量是其两倍。我们引入了广泛的新的评估套件,将多语言评估的最新水平扩展到 99 种语言——包括判别式和生成式任务、人工评估以及模拟胜率,这些评估既涵盖保持任务又涵盖分布内性能。此外,我们对最优微调混合物组成、数据剪枝以及模型的毒性、偏见和安全性进行了详细调查。我们在 https://hf.co/CohereForAI/aya-101 开源了我们的指令数据集和模型。

Recent breakthroughs in large language models (LLMs) have centered around a handful of data-rich languages. What does it take to broaden access to breakthroughs beyond first-class citizen languages? Our work introduces Aya, a massively multilingual generative language model that follows instructions in 101 languages of which over 50% are considered as lower-resourced. Aya outperforms mT0 and BLOOMZ on the majority of tasks while covering double the number of languages. We introduce extensive new evaluation suites that broaden the state-of-art for multilingual eval across 99 languages -- including discriminative and generative tasks, human evaluation, and simulated win rates that cover both held-out tasks and in-distribution performance. Furthermore, we conduct detailed investigations on the optimal finetuning mixture composition, data pruning, as well as the toxicity, bias, and safety of our models. We open-source our instruction datasets and our model at https://hf.co/CohereForAI/aya-101

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 30)

阅读逐段中英对照全文 →