HotpotQA:一个多样化、可解释的多跳问答数据集

HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering

杨植麟 Zhilin Yang · Carnegie Mellon University / Stanford / MILA · 2018-09-25 · arXiv:1809.09600 ↗ · 被引 4817

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

现有的问答数据集无法训练问答系统进行复杂推理并提供答案解释。我们引入了 HotpotQA,这是一个包含 11.3 万个基于维基百科的问答对的新数据集,具有四个关键特征:(1)问题需要查找并推理多个支持文档才能回答;(2)问题多样化,不受任何现有知识库或知识模式的约束;(3)我们提供推理所需的句子级支持事实,使问答系统能够在强监督下进行推理并解释预测;(4)我们提供了一种新类型的事实比较问题,以测试问答系统提取相关事实并进行必要比较的能力。我们表明,HotpotQA 对最新的问答系统具有挑战性,支持事实使模型能够提高性能并做出可解释的预测。

Existing question answering (QA) datasets fail to train QA systems to perform complex reasoning and provide explanations for answers. We introduce HotpotQA, a new dataset with 113k Wikipedia-based question-answer pairs with four key features: (1) the questions require finding and reasoning over multiple supporting documents to answer; (2) the questions are diverse and not constrained to any pre-existing knowledge bases or knowledge schemas; (3) we provide sentence-level supporting facts required for reasoning, allowing QA systems to reason with strong supervision and explain the predictions; (4) we offer a new type of factoid comparison questions to test QA systems' ability to extract relevant facts and perform necessary comparison. We show that HotpotQA is challenging for the latest QA systems, and the supporting facts enable models to improve performance and make explainable predictions.

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 12)

阅读逐段中英对照全文 →