Who Wins the Conflict? Mechanistic Interpretability of Text Bias in Audio LLMs
About
While Audio Large Language Models (Audio LLMs) excel at multimodal understanding, they suffer from text dominance, a bias where models blindly favor text over acoustic evidence, causing hallucinations. However, the internal mechanisms underlying how these models behave when audio and textual inputs contradict each other remain unexplored. In this work, we present the first mechanistic analysis of this phenomenon by tracing the propagation of internal representations across layers. Our investigation reveals three key findings: (i) text dominance is systematically and empirically across models; (ii) while text and audio rely on functionally distinct pathways, they ultimately converge into a shared semantic space in late layers; and (iii) the text pathway does not erase audio information, but rather actively suppresses intact audio representations. Building on these insights, we leverage back-patching, a training-free intervention that routes late-layer audio activations back into earlier layers. This amplifies the audio representations, enabling them to overcome textual suppression. Our evaluation shows that back-patching consistently reduces text dominance, paving the way for mechanistic multimodal alignment under conflict.
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Patched audio accuracy evaluation under modality conflict | Natural voice counterfactual audio samples | Adj. Swap Patched Acc59 | 18 | |
| Behavioral Accuracy | Natural Voice and Free Generation dataset | Text-Affinity Accuracy (Abase -> TCF)59.6 | 16 | |
| Modality Conflict Resolution | ALME | -- | 8 | |
| Behavioral accuracy evaluation | Synthetic TTS AR | Text-Affinity Accuracy (Abase -> TCF)70.8 | 2 | |
| Behavioral accuracy evaluation | Synthetic TTS DE | Text-Affinity Accuracy (Abase -> TCF)74 | 2 | |
| Behavioral accuracy evaluation | Synthetic TTS EN | Text-Affinity Accuracy (Abase -> TCF)0.74 | 2 | |
| Behavioral accuracy evaluation | Synthetic TTS FR | Text-Affinity Accuracy (Abase → TCF)71.6 | 2 | |
| Behavioral accuracy evaluation | Synthetic TTS IT | Text-Affinity Accuracy (Abase -> TCF)68.4 | 2 | |
| Behavioral accuracy evaluation | Synthetic TTS JA | Text-Affinity Accuracy (Abase -> TCF)66 | 2 | |
| Behavioral accuracy evaluation | Synthetic TTS PT | Text-Affinity Accuracy (Abase -> TCF)72 | 2 |