Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Who Wins the Conflict? Mechanistic Interpretability of Text Bias in Audio LLMs

About

While Audio Large Language Models (Audio LLMs) excel at multimodal understanding, they suffer from text dominance, a bias where models blindly favor text over acoustic evidence, causing hallucinations. However, the internal mechanisms underlying how these models behave when audio and textual inputs contradict each other remain unexplored. In this work, we present the first mechanistic analysis of this phenomenon by tracing the propagation of internal representations across layers. Our investigation reveals three key findings: (i) text dominance is systematically and empirically across models; (ii) while text and audio rely on functionally distinct pathways, they ultimately converge into a shared semantic space in late layers; and (iii) the text pathway does not erase audio information, but rather actively suppresses intact audio representations. Building on these insights, we leverage back-patching, a training-free intervention that routes late-layer audio activations back into earlier layers. This amplifies the audio representations, enabling them to overcome textual suppression. Our evaluation shows that back-patching consistently reduces text dominance, paving the way for mechanistic multimodal alignment under conflict.

Hyebin Cho, Suho Yoo, Jaehyuk Jang, Changick Kim, Joon Son Chung• 2026

Related benchmarks

TaskDatasetResultRank
Patched audio accuracy evaluation under modality conflictNatural voice counterfactual audio samples
Adj. Swap Patched Acc59
18
Behavioral AccuracyNatural Voice and Free Generation dataset
Text-Affinity Accuracy (Abase -> TCF)59.6
16
Modality Conflict ResolutionALME--
8
Behavioral accuracy evaluationSynthetic TTS AR
Text-Affinity Accuracy (Abase -> TCF)70.8
2
Behavioral accuracy evaluationSynthetic TTS DE
Text-Affinity Accuracy (Abase -> TCF)74
2
Behavioral accuracy evaluationSynthetic TTS EN
Text-Affinity Accuracy (Abase -> TCF)0.74
2
Behavioral accuracy evaluationSynthetic TTS FR
Text-Affinity Accuracy (Abase → TCF)71.6
2
Behavioral accuracy evaluationSynthetic TTS IT
Text-Affinity Accuracy (Abase -> TCF)68.4
2
Behavioral accuracy evaluationSynthetic TTS JA
Text-Affinity Accuracy (Abase -> TCF)66
2
Behavioral accuracy evaluationSynthetic TTS PT
Text-Affinity Accuracy (Abase -> TCF)72
2
Showing 10 of 11 rows

Other info

Follow for update