Insert Anything: Image Insertion via In-Context Editing in DiT

About

This work presents Insert Anything, a unified framework for reference-based image insertion that seamlessly integrates objects from reference images into target scenes under flexible, user-specified control guidance. Instead of training separate models for individual tasks, our approach is trained once on our new AnyInsertion dataset--comprising 120K prompt-image pairs covering diverse tasks such as person, object, and garment insertion--and effortlessly generalizes to a wide range of insertion scenarios. Such a challenging setting requires capturing both identity features and fine-grained details, while allowing versatile local adaptations in style, color, and texture. To this end, we propose to leverage the multimodal attention of the Diffusion Transformer (DiT) to support both mask- and text-guided editing. Furthermore, we introduce an in-context editing mechanism that treats the reference image as contextual information, employing two prompting strategies to harmonize the inserted elements with the target scene while faithfully preserving their distinctive features. Extensive experiments on AnyInsertion, DreamBooth, and VTON-HD benchmarks demonstrate that our method consistently outperforms existing alternatives, underscoring its great potential in real-world applications such as creative content generation, virtual try-on, and scene composition.

Wensong Song, Hong Jiang, Zongxing Yang, Ruijie Quan, Yi Yang• 2025

Related benchmarks

Task	Dataset	Result
Anomaly Localization	MVTec AD	Pixel AUROC97.9	543
Anomaly Detection	MVTec AD	Image AUROC84.97	92
Instructive image editing	MagicBrush (test)	CLIP Image0.8722	53
Cross-domain object insertion	AIComposer	CLIP-I84.87	21
Anomaly Detection	MVTec AD	Img AUROC95.3	16
Cross-domain compositing	TF-ICON & AIComposer Pooled	Identity Score2.33	15
Cross-domain compositing	TF-ICON	LPIPS0.7281	15
Reference-Guided Image Editing	UniEdit (test)	DINO-I56.22	14
Image Compositing	TF-ICON (test)	LPIPS0.7281	13
Object Compositing	DreamBooth (test)	Fidelity Score18.4	10

Showing 10 of 26 rows

Other info

Follow for update

@wizwand_team Discord