SegGPT: 在上下文中分割一切

SegGPT: Segmenting Everything In Context

王鑫龙 Xinlong Wang · BAAI · 2023-04-06 · arXiv:2304.03284 ↗ · 被引 275

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

我们提出了 SegGPT,一个用于在上下文中分割一切的通用模型。我们将各种分割任务统一到一个通用的上下文学习框架中,通过将不同类型的分割数据转换为相同格式的图像来适应它们。SegGPT 的训练被表述为一个上下文着色问题,每个数据样本具有随机颜色映射。目标是根据上下文完成不同的任务,而不是依赖特定的颜色。训练后,SegGPT 可以通过上下文推理在图像或视频中执行任意分割任务,例如对象实例、材质、部分、轮廓和文本。SegGPT 在广泛的任务上进行了评估,包括少样本语义分割、视频对象分割、语义分割和全景分割。我们的结果在定性和定量上都显示出在分割域内和域外目标方面的强大能力。

We present SegGPT, a generalist model for segmenting everything in context. We unify various segmentation tasks into a generalist in-context learning framework that accommodates different kinds of segmentation data by transforming them into the same format of images. The training of SegGPT is formulated as an in-context coloring problem with random color mapping for each data sample. The objective is to accomplish diverse tasks according to the context, rather than relying on specific colors. After training, SegGPT can perform arbitrary segmentation tasks in images or videos via in-context inference, such as object instance, stuff, part, contour, and text. SegGPT is evaluated on a broad range of tasks, including few-shot semantic segmentation, video object segmentation, semantic segmentation, and panoptic segmentation. Our results show strong capabilities in segmenting in-domain and out-of-domain targets, either qualitatively or quantitatively.

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 18)

阅读逐段中英对照全文 →