Microsoft COCO:上下文中的常见物体

Microsoft COCO: Common Objects in Context

林宗毅 Tsung-Yi Lin · Microsoft · 2014-05-01 · arXiv:1405.0312 ↗ · 被引 53675

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

我们提出了一个新数据集,旨在通过将物体识别问题置于场景理解这一更广泛问题的背景下,推动物体识别技术的发展。为此,我们收集了包含常见物体在其自然环境中出现的复杂日常场景图像。物体使用逐实例分割进行标注,以帮助精确定位物体。我们的数据集包含 91 种物体类型的照片,这些类型对于 4 岁儿童来说易于识别。数据集共有 328k 张图像中的 250 万个标注实例,其创建过程通过新颖的用户界面(用于类别检测、实例定位和实例分割)吸引了大量众包工人的参与。我们提供了与 PASCAL、ImageNet 和 SUN 数据集的详细统计分析比较。最后,我们使用可变形部件模型提供了边界框和分割检测结果的基线性能分析。

We present a new dataset with the goal of advancing the state-of-the-art in object recognition by placing the question of object recognition in the context of the broader question of scene understanding. This is achieved by gathering images of complex everyday scenes containing common objects in their natural context. Objects are labeled using per-instance segmentations to aid in precise object localization. Our dataset contains photos of 91 objects types that would be easily recognizable by a 4 year old. With a total of 2.5 million labeled instances in 328k images, the creation of our dataset drew upon extensive crowd worker involvement via novel user interfaces for category detection, instance spotting and instance segmentation. We present a detailed statistical analysis of the dataset in comparison to PASCAL, ImageNet, and SUN. Finally, we provide baseline performance analysis for bounding box and segmentation detection results using a Deformable Parts Model.

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 19)

阅读逐段中英对照全文 →