Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

ESC: Emotional Self-Correction for Reliable Vision-Language Models

About

Vision-language models (VLMs) have achieved strong performance across diverse multimodal tasks, yet they remain vulnerable to unreliable reasoning. Existing self-correction methods mitigate these issues but typically rely on post-training or carefully engineered feedback, incurring high computational cost. In this work, we revisit this challenge through the lens of emotional cues, asking whether they can activate latent self-correction behaviors in VLMs without additional training. \textbf{We find that emotional signals serve as an effective trigger for self-correction, encouraging more cautious and reflective reasoning}. Motivated by this finding, we propose \escabstract (\textbf{\underline{E}}motional \textbf{\underline{S}}elf-\textbf{\underline{C}}orrection), a training-free self-correction framework. ESC introduces an external verifier that detects potentially incorrect initial responses and injects emotional feedback to encourage model to reflect, and produce a better revised response without additional training. Extensive experiments across safety, hallucination, vision-centric perception, and multimodal reasoning benchmarks show that ESC consistently improves reliability while preserving overall model utility. These results suggest that emotion can function not only as an ability to be recognized, but also as a practical control signal for scalable self-correction in VLMs. \textbf{We therefore believe that ESC provides a strong foundation for a new reliable human-like, emotion-integrated research direction.} Our project is publicly available at \textcolor{red}{https://genai4e.github.io/ESC/}.

Tien-Huy Nguyen, Minh-Nhat Nguyen, Nguyen Nhat Huy, Hung Viet Nguyen, Huy Nguyen Minh Nhat, Thanh-Huy Nguyen, Cuong Tuan Nguyen, Hoang M. Le, Dat Nguyen, Phat Kim Huynh, Min Xu, Ulas Bagci• 2026

Related benchmarks

TaskDatasetResultRank
Multimodal ReasoningMM-Vet
MM-Vet Score40.05
551
Multimodal ReasoningMMMU
Accuracy40.02
220
Object Hallucination EvaluationPOPE Adversarial
Accuracy84.53
174
Hallucination EvaluationHallusionBench--
153
Multimodal ReasoningMMStar
Accuracy56.07
102
Object Hallucination EvaluationPOPE Popular
Accuracy87.4
100
Multimodal ReasoningMathVista
Accuracy56.1
50
Hallucination EvaluationPOPE (Random)
Accuracy89.57
11
Multimodal ReasoningAI2D
Accuracy62.53
4
Vision-Centric PerceptionMMVP 84
Pair Accuracy45.33
4
Showing 10 of 13 rows

Other info

Follow for update