Seedream 4.0: Toward Next-generation Multimodal Image Generation
打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→我们介绍了 Seedream 4.0,一个高效且高性能的多模态图像生成系统,它将文本到图像(T2I)合成、图像编辑和多图像组合统一在单一框架中。我们开发了一个高效的扩散变换器,并配备了一个强大的 VAE,该 VAE 还能显著减少图像 token 的数量。这使得模型训练高效,并能快速生成原生高分辨率图像(例如 1K-4K)。Seedream 4.0 在数十亿涵盖多种分类和知识中心概念的文本-图像对上进行了预训练。通过跨数百个垂直场景的全面数据收集,结合优化策略,确保了稳定的大规模训练和强大的泛化能力。通过整合一个精心微调的 VLM 模型,我们进行了多模态后训练,联合训练 T2I 和图像编辑任务。为了加速推理,我们集成了对抗性蒸馏、分布匹配、量化以及推测解码。生成一张 2K 图像的推理时间可达 1.8 秒(无需 LLM/VLM 作为 PE 模型)。全面评估显示,Seedream 4.0 在 T2I 和多模态图像编辑上均达到了最先进的结果。特别是,它在复杂任务中展现出卓越的多模态能力,包括精确图像编辑和上下文推理,还支持多图像参考,并能生成多个输出图像。这扩展了传统的 T2I 系统,使其成为更具交互性和多维度的创意工具,推动了生成式 AI 在创意和专业应用领域的边界。我们进一步将模型和数据扩展为 Seedream 4.5。Seedream 4.0 和 Seedream 4.5 可在火山引擎上访问:https://www.volcengine.com/experience/ark?launch=seedream。
We introduce Seedream 4.0, an efficient and high-performance multimodal image generation system that unifies text-to-image (T2I) synthesis, image editing, and multi-image composition within a single framework. We develop a highly efficient diffusion transformer with a powerful VAE which also can reduce the number of image tokens considerably. This allows for efficient training of our model, and enables it to fast generate native high-resolution images (e.g., 1K-4K). Seedream 4.0 is pretrained on billions of text-image pairs spanning diverse taxonomies and knowledge-centric concepts. Comprehensive data collection across hundreds of vertical scenarios, coupled with optimized strategies, ensures stable and large-scale training, with strong generalization. By incorporating a carefully fine-tuned VLM model, we perform multi-modal post-training for training both T2I and image editing tasks jointly. For inference acceleration, we integrate adversarial distillation, distribution matching, and quantization, as well as speculative decoding. It achieves an inference time of up to 1.8 seconds for generating a 2K image (without a LLM/VLM as PE model). Comprehensive evaluations reveal that Seedream 4.0 can achieve state-of-the-art results on both T2I and multimodal image editing. In particular, it demonstrates exceptional multimodal capabilities in complex tasks, including precise image editing and in-context reasoning, and also allows for multi-image reference, and can generate multiple output images. This extends traditional T2I systems into an more interactive and multidimensional creative tool, pushing the boundary of generative AI for both creativity and professional applications. We further scale our model and data as Seedream 4.5. Seedream 4.0 and Seedream 4.5 are accessible on Volcano Engine https://www.volcengine.com/experience/ark?launch=seedream.