ByteDance's DanceGRPO Framework Boosts Visual Generation Capabilities
ByteDance's AI research division has developed DanceGRPO, a framework that enhances post-training model capabilities in visual generation tasks. The team adapted Group Relative Policy Optimization (GRPO), originally designed for large language models, to work with diffusion and rectified flow models.
The innovation involves reworking the sampling process to make GRPO's group-based relative scoring compatible with visual pipelines. This allows DanceGRPO to support multiple foundational models, including Stable Diffusion, FLUX, HunyuanVideo, and SkyReels-I2V, while giving researchers flexibility in defining 'good' output.
Benchmark improvements show enhancements of up to 181% on metrics like HPS-v2.1 and CLIP Score, which measure image alignment with text prompts and human preferences. A follow-on project called BranchGRPO reported alignment score improvements of up to 16% while cutting training time by 55%