Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

FADE: Mitigating Hallucinations by Reducing Language-Prior Dominance in Large Vision-Language Models

About

Despite the impressive capabilities of Large Vision-Language Models (LVLMs), they remain susceptible to hallucination, generating content inconsistent with the input image. Recent studies attribute this to the dominance of language priors over visual inputs and employ contrastive decoding methods to mitigate this dominance, but the mechanistic origin remains unexplored. We investigate the information flow through each transformer layer and find that attention modules consistently aggregate visual evidence, while FFN modules at critical layers act as the source of language priors. These priors can override visual evidence, causing correct predictions in intermediate layers to drift toward incorrect outputs. Based on this insight, we propose FADE (FFN Attenuation for DEcoding), a training-free method that attenuates FFN outputs to reduce language-prior dominance. Evaluations on POPE, CHAIR, and MME benchmarks across LLaVA-1.5, mPLUG-Owl2, and InstructBLIP show that FADE effectively mitigates hallucinations while preserving inference efficiency.

Yichen Guo, Kai Tang, Fenglai Lin, Yiding Sun, Dongxu Zhang, Wenya Wang, Lin William Cong, Shanghang Zhang• 2026

Related benchmarks

TaskDatasetResultRank
Object Hallucination EvaluationPOPE--
2056
Multimodal EvaluationMME--
902
Multimodal Model EvaluationMMBench--
265
Multimodal ReasoningMMBench--
180
Object Hallucination EvaluationCHAIR--
174
Multimodal EvaluationMMBench--
146
Caption Hallucination EvaluationCHAIR
CS Score36.6
122
Object Hallucination EvaluationPOPE MSCOCO
F1 Score88.3
114
Object Hallucination EvaluationPOPE GQA
Accuracy84.1
86
Object Hallucination EvaluationPOPE A-OKVQA
Accuracy91.2
83
Showing 10 of 20 rows

Other info

Follow for update