We present a new dataset with the goal of advancing the state-of-the-art in object recognition by placing the question of object recognition in the context of the broader question of scene understanding. This is achieved by gathering images of complex everyday scenes containing common objects in their natural context. Objects are labeled using per-instance segmentations to aid in precise object localization. Our dataset contains photos of 91 objects types that would be easily recognizable by a 4 year old. With a total of 2.5 million labeled instances in 328k images, the creation of our dataset drew upon extensive crowd worker involvement via novel user interfaces for category detection, instance spotting and instance segmentation. We present a detailed statistical analysis of the dataset in comparison to PASCAL, ImageNet, and SUN. Finally, we provide baseline performance analysis for bounding box and segmentation detection results using a Deformable Parts Model.
核心贡献 · Key contributions
提出 MS COCO 数据集,包含 91 个物体类别、328k 张图像中的 250 万个标注实例。 Introduces MS COCO dataset with 91 object categories, 2.5M instances in 328k images.
聚焦非典型视角图像,强调上下文关系与精确的实例分割。 Focuses on non-iconic images with contextual relationships and precise instance segmentation.
创新标注流程,采用分层标注与众包提高效率。 Novel annotation pipeline using hierarchical labeling and crowdsourcing for efficiency.
与 PASCAL、ImageNet 和 SUN 数据集进行详细统计对比。 Provides detailed statistical comparison with PASCAL, ImageNet, and SUN datasets.
使用可变形部件模型提供检测与分割的基线结果。 Baseline detection and segmentation results using Deformable Parts Model.
展示数据集难度及跨数据集泛化的优势。 Demonstrates dataset difficulty and cross-dataset generalization benefits.
局限 · Limitations
仅包含“物体”类别,不包括“材质”如天空或草地。 Only includes 'thing' categories, not 'stuff' like sky or grass.