Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Scaling Multi-Reference Image Generation with Dynamic Reward Optimization

About

While personalized image generation has achieved remarkable progress, multi-reference image generation (MRIG) remains a challenging task. Most existing benchmarks fail to adequately evaluate complex MRIG scenarios, hindering further progress in this area. To better assess model performance on complex MRIG tasks, we introduce OmniRef-Bench, a benchmark that covers complex combinations of reference image types and a large number of reference images. Evaluations on OmniRef-Bench show that mainstream open-source models struggle in complex MRIG scenarios, and their performance deteriorates significantly as the number of mixed-type reference images increases. To address this issue, we propose DyRef, a two-stage training framework. In the first stage, supervised fine-tuning equips the model with the basic capability to handle complex MRIG tasks. In the second stage, we introduce Difficulty-aware Advantage Reweighting (DAR) and Discriminative Reward Scaling (DRS). DAR dynamically adjusts the optimization objective to improve performance when handling a large number of mixed-type reference images. DRS enlarges intra-group reward differences for more effective policy optimization. Experiments demonstrate that DyRef significantly improves the performance of open-source models on OmniRef-Bench and single-image editing benchmarks, demonstrating the effectiveness and generalization capability of our approach.

Wenwang Huang, Yusen Fu, Junjie Wang, Mengfei Huang, Yulin Li, Gan Liu, Jing Cai, Yancheng He, Zhuotao Tian• 2026

Related benchmarks

TaskDatasetResultRank
Subject-consistent image generationOmniContext
Fidelity (Single, Character)8.59
15
Single-image editingImgEdit Benchmark
Extract Score4.07
9
Multi-reference image generationOmniRef-Bench
Objective Subject Score0.63
9
Multi-reference image generationOmniRef-Bench 1.0 (test)
Score8.38
8
Multi-reference image generationUser Study 50 samples
Subjective Score8.38
8
Multi-reference image generationDreambench++
CP (Animal)68
7
Image EditingDreambench++
DINOv1 Score0.6
5
Multi-reference image generationMultiBanana
Fidelity (Single Object)6.34
4
Multi-reference image generationMultiBanana
Text Quality Score3.82
2
Multi-reference image generationMultiRef
BBox Score8.71
2
Showing 10 of 10 rows

Other info

Follow for update