We present evaluation results for FLUX.1 Kontext, a generative flow matching model that unifies image generation and editing. The model generates novel output views by incorporating semantic context from text and image inputs. Using a simple sequence concatenation approach, FLUX.1 Kontext handles both local editing and generative in-context tasks within a single unified architecture. Compared to current editing models that exhibit degradation in character consistency and stability across multiple turns, we observe that FLUX.1 Kontext improved preservation of objects and characters, leading to greater robustness in iterative workflows. The model achieves competitive performance with current state-of-the-art systems while delivering significantly faster generation times, enabling interactive applications and rapid prototyping workflows. To validate these improvements, we introduce KontextBench, a comprehensive benchmark with 1026 image-prompt pairs covering five task categories: local editing, global editing, character reference, style reference and text editing. Detailed evaluations show the superior performance of FLUX.1 Kontext in terms of both single-turn quality and multi-turn consistency, setting new standards for unified image processing models.
核心贡献 · Key contributions
通过简单的序列拼接,在单一流匹配架构中统一图像生成与编辑。 Unifies image generation and editing in a single flow matching architecture via simple sequence concatenation.
在多次迭代编辑中实现卓越的角色与物体一致性,减少视觉漂移。 Achieves superior character and object consistency across multiple iterative edits, reducing visual drift.
显著加快生成速度(1024×1024 分辨率下 3-5 秒),支持交互式应用。 Delivers significantly faster generation times (3–5 seconds at 1024×1024) enabling interactive applications.
提出 KontextBench,包含 1026 个真实世界图像-提示对,覆盖五个任务类别。 Introduces KontextBench, a comprehensive benchmark with 1026 real-world image-prompt pairs across five task categories.
采用潜在对抗扩散蒸馏(LADD)减少采样步数同时提升输出质量。 Employs latent adversarial diffusion distillation (LADD) to reduce sampling steps while improving output quality.
在角色一致性和速度上优于 GPT-Image-1 等竞品模型,最高快一个数量级。 Outperforms competing models like GPT-Image-1 in character consistency and speed by up to an order of magnitude.
局限 · Limitations
过度多轮编辑可能引入视觉伪影,降低图像质量。 Excessive multi-turn editing can introduce visual artifacts that degrade image quality.
模型偶尔未能准确遵循指令,忽略特定提示要求。 Model occasionally fails to follow instructions accurately, ignoring specific prompt requirements.
蒸馏过程可能引入视觉伪影,影响输出保真度。 Distillation process can introduce visual artifacts impacting output fidelity.
当前仅支持单张上下文图像条件;扩展到多张图像是未来工作。 Currently limited to single context image conditioning; extension to multiple images is future work.
评估聚焦于静态图像;视频域编辑尚未涉及。 Evaluation focuses on static images; video domain editing is not yet addressed.