FLUX.1 Kontext:潜空间上下文图像生成与编辑的流匹配方法

FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space

罗宾·罗姆巴赫 Robin Rombach · Black Forest Labs · 2025-06-17 · arXiv:2506.15742 ↗ · 被引 864

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

我们展示了 FLUX.1 Kontext 的评估结果,这是一种生成式流匹配模型,统一了图像生成与编辑。该模型通过整合文本和图像输入的语义上下文,生成新颖的输出视图。采用简单的序列拼接方法,FLUX.1 Kontext 在单一统一架构中处理局部编辑和生成式上下文任务。与当前在多轮编辑中表现出角色一致性和稳定性下降的模型相比,我们观察到 FLUX.1 Kontext 改进了对象和角色的保持,从而在迭代工作流中具有更强的鲁棒性。该模型在实现显著更快的生成时间的同时,达到了与当前最先进系统相竞争的性能,支持交互式应用和快速原型设计工作流。为验证这些改进,我们引入了 KontextBench,一个包含 1026 个图像-提示对的综合基准,涵盖五个任务类别:局部编辑、全局编辑、角色参考、风格参考和文本编辑。详细评估显示,FLUX.1 Kontext 在单轮质量和多轮一致性方面均表现出优越性能,为统一图像处理模型设立了新标准。

We present evaluation results for FLUX.1 Kontext, a generative flow matching model that unifies image generation and editing. The model generates novel output views by incorporating semantic context from text and image inputs. Using a simple sequence concatenation approach, FLUX.1 Kontext handles both local editing and generative in-context tasks within a single unified architecture. Compared to current editing models that exhibit degradation in character consistency and stability across multiple turns, we observe that FLUX.1 Kontext improved preservation of objects and characters, leading to greater robustness in iterative workflows. The model achieves competitive performance with current state-of-the-art systems while delivering significantly faster generation times, enabling interactive applications and rapid prototyping workflows. To validate these improvements, we introduce KontextBench, a comprehensive benchmark with 1026 image-prompt pairs covering five task categories: local editing, global editing, character reference, style reference and text editing. Detailed evaluations show the superior performance of FLUX.1 Kontext in terms of both single-turn quality and multi-turn consistency, setting new standards for unified image processing models.

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 13)

阅读逐段中英对照全文 →