Locate-then-Sparsify: Attribution Guided Sparse Strategy for Visual Hallucination Mitigation

About

Despite the significant advancements in Large Vision-Language Models (LVLMs), their tendency to generate hallucinations undermines reliability and restricts broader practical deployment. Among the hallucination mitigation methods, feature steering emerges as a promising approach that reduces erroneous outputs in LVLMs without increasing inference costs. However, current methods apply uniform feature steering across all layers. This heuristic strategy ignores inter-layer differences, potentially disrupting layers unrelated to hallucinations and ultimately leading to performance degradation on general tasks. In this paper, we propose Locate-Then-Sparsify for Feature Steering (LTS-FS), a plug-and-play framework which controls the steering intensity according to the hallucination relevance of each layer. We first construct a dataset comprising token-level and sentence-level hallucination cases. Based on this dataset, we introduce an attribution method based on causal interventions to quantify the hallucination relevance of each layer. With the attribution scores across layers, we propose a layerwise strategy that converts these scores into feature steering intensities for individual layers, enabling more precise adjustments specifically on hallucination-relevant layers. Extensive experiments across multiple LVLMs and benchmarks demonstrate that LTS-FS effectively mitigates hallucination while preserving strong performance. Codes are available at https://github.com/huttersadan/LTS-FS.

Tiantian Dang, Chao Bi, Shufan Shen, Jinzhe Liu, Qingming Huang, Shuhui Wang• 2026

Related benchmarks

Task	Dataset	Result
Object Hallucination Evaluation	POPE	Accuracy79.92	2019
Hallucination Evaluation	CHAIR	CHAIR_s46.8	393
Object Hallucination	POPE Popular	F1 Score83.58	372
Object Hallucination	POPE (Random)	F1 Score87.64	324
Object Hallucination Evaluation	POPE Adversarial	--	159
Caption Hallucination Evaluation	CHAIR	CS Score46.8	44
Hallucination Evaluation	MSCOCO	CS Score46.8	21
Object Hallucination Evaluation	POPE GQA	Accuracy77.15	20
Multimodal Assistant Evaluation	LLaVA-Bench GPT-4V-aided (full)	Accuracy6.96	6
Generative Capability Evaluation	CLAIR	CLAIR Details Score6.23	4

Showing 10 of 10 rows

Other info

Follow for update

@wizwand_team Discord