LEGION: Learning to Ground and Explain for Synthetic Image Detection

About

The rapid advancements in generative technology have emerged as a double-edged sword. While offering powerful tools that enhance convenience, they also pose significant social concerns. As defenders, current synthetic image detection methods often lack artifact-level textual interpretability and are overly focused on image manipulation detection, and current datasets usually suffer from outdated generators and a lack of fine-grained annotations. In this paper, we introduce SynthScars, a high-quality and diverse dataset consisting of 12,236 fully synthetic images with human-expert annotations. It features 4 distinct image content types, 3 categories of artifacts, and fine-grained annotations covering pixel-level segmentation, detailed textual explanations, and artifact category labels. Furthermore, we propose LEGION (LEarning to Ground and explain for Synthetic Image detectiON), a multimodal large language model (MLLM)-based image forgery analysis framework that integrates artifact detection, segmentation, and explanation. Building upon this capability, we further explore LEGION as a controller, integrating it into image refinement pipelines to guide the generation of higher-quality and more realistic images. Extensive experiments show that LEGION outperforms existing methods across multiple benchmarks, particularly surpassing the second-best traditional expert on SynthScars by 3.31% in mIoU and 7.75% in F1 score. Moreover, the refined images generated under its guidance exhibit stronger alignment with human preferences. The code, model, and dataset will be released.

Hengrui Kang, Siwei Wen, Zichen Wen, Junyan Ye, Weijia Li, Peilin Feng, Baichuan Zhou, Bin Wang, Dahua Lin, Linfeng Zhang, Conghui He• 2025

Related benchmarks

Task	Dataset	Result
Image Forgery Detection	Trace (test)	Accuracy65.4	18
Visual Reasoning	Trace (test)	BLEU-10.102	17
Artifact Correction	SynthScars (test)	GPT-assisted Structure Score0.35	11
Generative image detection	FakeClue (test)	Overall Accuracy25.4	11
Artifact Localization	SynthScars (test)	mIoU0.106	10
Artifact Localization	LOKI (test)	mIoU0.1	10
Artifact Localization	ArtiBench (test)	mIoU0.062	10
Artifact Localization	RichHF (test)	mIoU6.7	10
Forgery Grounding	Trace (test)	IoU6.1	10
Artifact Explanation	SynthScars (test)	ROUGE Score24.7	8

Showing 10 of 20 rows

Other info

Follow for update

@wizwand_team Discord