Divide & Bind Your Attention for Improved Generative Semantic Nursing

About

Emerging large-scale text-to-image generative models, e.g., Stable Diffusion (SD), have exhibited overwhelming results with high fidelity. Despite the magnificent progress, current state-of-the-art models still struggle to generate images fully adhering to the input prompt. Prior work, Attend & Excite, has introduced the concept of Generative Semantic Nursing (GSN), aiming to optimize cross-attention during inference time to better incorporate the semantics. It demonstrates promising results in generating simple prompts, e.g., "a cat and a dog". However, its efficacy declines when dealing with more complex prompts, and it does not explicitly address the problem of improper attribute binding. To address the challenges posed by complex prompts or scenarios involving multiple entities and to achieve improved attribute binding, we propose Divide & Bind. We introduce two novel loss objectives for GSN: a novel attendance loss and a binding loss. Our approach stands out in its ability to faithfully synthesize desired objects with improved attribute alignment from complex prompts and exhibits superior performance across multiple evaluation benchmarks.

Yumeng Li, Margret Keuper, Dan Zhang, Anna Khoreva• 2023

Related benchmarks

Task	Dataset	Result
Text-to-Image Generation	Multi-subject prompts (test)	CLIP I-T0.3489	10
Text-to-Image Generation	SD 3.5	Win Rate (%)50	10
Text-to-Image Generation	FLUX.1	Win Rate50	10
Numerical Reasoning	HRS	Precision77.95	8
Numerical Reasoning	NSR-1K	Precision84.99	8
Coarse-grained attribute binding	Author's Benchmark 1.0 (test)	VQAScore0.372	8
Coarse-grained attribute binding	User Study 10 prompts (test)	User Preference Frequency3.2	8
Spatial Reasoning	HRS	Accuracy13.15	8
Spatial Reasoning	NSR-1K	Accuracy24	8
Text-to-Image Synthesis	User study 20 questions (test)	User Preference Rate6.67	7

Showing 10 of 18 rows

Other info

Follow for update

@wizwand_team Discord