Seedream 3.0 技术报告

Seedream 3.0 Technical Report

高宇 Yu Gao · ByteDance Seed · 2025-04-15 · arXiv:2504.11346 ↗ · 被引 134

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

我们提出了 Seedream 3.0,一个高性能的中英双语图像生成基础模型。我们开发了多项技术改进,以解决 Seedream 2.0 中存在的挑战,包括复杂提示的对齐、精细排版生成、次优的视觉美学和保真度以及有限的图像分辨率。具体来说,Seedream 3.0 的进步源于从数据构建到模型部署的整个流程的改进。在数据层面,我们通过缺陷感知训练范式和双轴协同数据采样框架将数据集扩大了一倍。此外,我们在预训练阶段采用了多种有效技术,如混合分辨率训练、跨模态 RoPE、表示对齐损失和分辨率感知时间步采样。在后训练阶段,我们利用多样化美学描述进行 SFT,并采用基于 VLM 的奖励模型进行缩放,从而实现了与人类偏好良好对齐的输出。此外,Seedream 3.0 开创了一种新的加速范式。通过采用一致噪声期望和重要性感知时间步采样,我们在保持图像质量的同时实现了 4 到 8 倍的加速。Seedream 3.0 相比 Seedream 2.0 表现出显著改进:它增强了整体能力,特别是在复杂中文字符的文本渲染方面,这对专业排版生成至关重要。此外,它提供了原生高分辨率输出(高达 2K),能够生成具有高视觉质量的图像。

We present Seedream 3.0, a high-performance Chinese-English bilingual image generation foundation model. We develop several technical improvements to address existing challenges in Seedream 2.0, including alignment with complicated prompts, fine-grained typography generation, suboptimal visual aesthetics and fidelity, and limited image resolutions. Specifically, the advancements of Seedream 3.0 stem from improvements across the entire pipeline, from data construction to model deployment. At the data stratum, we double the dataset using a defect-aware training paradigm and a dual-axis collaborative data-sampling framework. Furthermore, we adopt several effective techniques such as mixed-resolution training, cross-modality RoPE, representation alignment loss, and resolution-aware timestep sampling in the pre-training phase. During the post-training stage, we utilize diversified aesthetic captions in SFT, and a VLM-based reward model with scaling, thereby achieving outputs that well align with human preferences. Furthermore, Seedream 3.0 pioneers a novel acceleration paradigm. By employing consistent noise expectation and importance-aware timestep sampling, we achieve a 4 to 8 times speedup while maintaining image quality. Seedream 3.0 demonstrates significant improvements over Seedream 2.0: it enhances overall capabilities, in particular for text-rendering in complicated Chinese characters which is important to professional typography generation. In addition, it provides native high-resolution output (up to 2K), allowing it to generate images with high visual quality.

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 13)

阅读逐段中英对照全文 →