Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Mitigating Hallucinations via Inter-Layer Consistency Aggregation in Large Vision-Language Models

About

Despite the impressive capabilities of Large Vision-Language Models (LVLMs), they remain susceptible to hallucinations, where generated content is inconsistent with the input image. Existing training-free hallucination mitigation methods often suffer from unstable performance and high sensitivity to hyperparameter settings, which limits their practicality and broader adoption. In this paper, we propose Decoding with Inter-layer Consistency via Layer Aggregation (DCLA), a training-free decoding mechanism that requires no retraining, fine-tuning, or access to external knowledge bases. Specifically, DCLA constructs a dynamic semantic reference by aggregating representations from previous layers and uses it to correct semantically deviated layers, thereby enforcing inter-layer consistency. Experiments across seven LVLMs and multiple benchmarks demonstrate the generality of DCLA: it surpasses standard decoding by 28.58 MME points on LLaVA1.5-7B and 42.6 MME points on Qwen2.5-VL, while improving POPE accuracy by 2.74 percentage points in the strongest setting.

Kai Tang, Jinhao You, Yichen Guo, Yiding Sun, Dongxu Zhang, Wenya Wang, Hanze Li, Tao Luo, Renyuan Li, Xiande Huang, Shanghang Zhang• 2025

Related benchmarks

TaskDatasetResultRank
Object Hallucination EvaluationPOPE--
2056
Multimodal EvaluationMME--
902
Multimodal Model EvaluationMMBench--
265
Multimodal ReasoningMMBench--
180
Multimodal EvaluationMMStar--
177
Object Hallucination EvaluationCHAIR--
174
Multimodal EvaluationMMBench
MMB^CN Score85.74
146
Visual Question AnsweringVizWiz (test)--
136
Caption Hallucination EvaluationCHAIR
CS Score37.4
122
Object Hallucination EvaluationPOPE GQA
Accuracy91
86
Showing 10 of 32 rows

Other info

Follow for update