DanceGRPO: Unleashing GRPO on Visual Generation
打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→生成式 AI 的最新进展彻底改变了视觉内容创作,但将模型输出与人类偏好对齐仍是一个关键挑战。虽然强化学习已成为微调生成模型的一种有前景的方法,但现有方法如 DDPO 和 DPOK 存在根本性局限——特别是在扩展到大规模和多样化提示集时无法保持稳定优化,严重限制了其实用性。本文提出 DanceGRPO 框架,通过创新性地将群体相对策略优化(GRPO)应用于视觉生成任务来解决这些局限。我们的关键洞察是,GRPO 固有的稳定机制使其能够克服先前基于强化学习方法在视觉生成中面临的优化挑战。DanceGRPO 取得了多项重要进展:首先,它在多种现代生成范式(包括扩散模型和整流流)中展示了持续稳定的策略优化。其次,在扩展到包含三个关键任务和四个基础模型的复杂真实场景时,它保持了稳健性能。第三,它在优化由五种不同奖励模型(评估图像/视频美学、文本-图像对齐、视频运动质量和二元反馈)捕获的多样化人类偏好方面表现出显著的多功能性。我们的综合实验表明,DanceGRPO 在多个既定基准(包括 HPS-v2.1、CLIP Score、VideoAlign 和 GenEval)上比基线方法高出高达 181%。我们的结果确立了 DanceGRPO 作为在视觉生成中扩展基于人类反馈的强化学习任务的稳健且多功能的解决方案,为协调强化学习和视觉合成提供了新见解。
Recent advances in generative AI have revolutionized visual content creation, yet aligning model outputs with human preferences remains a critical challenge. While Reinforcement Learning (RL) has emerged as a promising approach for fine-tuning generative models, existing methods like DDPO and DPOK face fundamental limitations - particularly their inability to maintain stable optimization when scaling to large and diverse prompt sets, severely restricting their practical utility. This paper presents DanceGRPO, a framework that addresses these limitations through an innovative adaptation of Group Relative Policy Optimization (GRPO) for visual generation tasks. Our key insight is that GRPO's inherent stability mechanisms uniquely position it to overcome the optimization challenges that plague prior RL-based approaches on visual generation. DanceGRPO establishes several significant advances: First, it demonstrates consistent and stable policy optimization across multiple modern generative paradigms, including both diffusion models and rectified flows. Second, it maintains robust performance when scaling to complex, real-world scenarios encompassing three key tasks and four foundation models. Third, it shows remarkable versatility in optimizing for diverse human preferences as captured by five distinct reward models assessing image/video aesthetics, text-image alignment, video motion quality, and binary feedback. Our comprehensive experiments reveal that DanceGRPO outperforms baseline methods by up to 181\% across multiple established benchmarks, including HPS-v2.1, CLIP Score, VideoAlign, and GenEval. Our results establish DanceGRPO as a robust and versatile solution for scaling Reinforcement Learning from Human Feedback (RLHF) tasks in visual generation, offering new insights into harmonizing reinforcement learning and visual synthesis.