Scaling Multi-Reference Image Generation with Dynamic Reward Optimization
About
While personalized image generation has achieved remarkable progress, multi-reference image generation (MRIG) remains a challenging task. Most existing benchmarks fail to adequately evaluate complex MRIG scenarios, hindering further progress in this area. To better assess model performance on complex MRIG tasks, we introduce OmniRef-Bench, a benchmark that covers complex combinations of reference image types and a large number of reference images. Evaluations on OmniRef-Bench show that mainstream open-source models struggle in complex MRIG scenarios, and their performance deteriorates significantly as the number of mixed-type reference images increases. To address this issue, we propose DyRef, a two-stage training framework. In the first stage, supervised fine-tuning equips the model with the basic capability to handle complex MRIG tasks. In the second stage, we introduce Difficulty-aware Advantage Reweighting (DAR) and Discriminative Reward Scaling (DRS). DAR dynamically adjusts the optimization objective to improve performance when handling a large number of mixed-type reference images. DRS enlarges intra-group reward differences for more effective policy optimization. Experiments demonstrate that DyRef significantly improves the performance of open-source models on OmniRef-Bench and single-image editing benchmarks, demonstrating the effectiveness and generalization capability of our approach.
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Subject-consistent image generation | OmniContext | Fidelity (Single, Character)8.59 | 15 | |
| Single-image editing | ImgEdit Benchmark | Extract Score4.07 | 9 | |
| Multi-reference image generation | OmniRef-Bench | Objective Subject Score0.63 | 9 | |
| Multi-reference image generation | OmniRef-Bench 1.0 (test) | Score8.38 | 8 | |
| Multi-reference image generation | User Study 50 samples | Subjective Score8.38 | 8 | |
| Multi-reference image generation | Dreambench++ | CP (Animal)68 | 7 | |
| Image Editing | Dreambench++ | DINOv1 Score0.6 | 5 | |
| Multi-reference image generation | MultiBanana | Fidelity (Single Object)6.34 | 4 | |
| Multi-reference image generation | MultiBanana | Text Quality Score3.82 | 2 | |
| Multi-reference image generation | MultiRef | BBox Score8.71 | 2 |