Omni-Perception Policy Optimization for Multimodal Emotion Reasoning
About
We find that current emotion-oriented Omni-MLLMs still lack reliable omni-modal perception: they (i) underutilize multimodal cues in their reasoning trajectories and (ii) exhibit unfaithful behavior, often hallucinating modality-specific statements from other modalities. Building on these insights, we propose OPPO (Omni-Perception Policy Optimization), a reinforcement learning framework that explicitly optimizes multimodal perception. First, an Omni-Perception Reward decomposes ground-truth reasoning into fine-grained visual, acoustic, and emotion cues and rewards trajectories that semantically recover these cues. Second, an Omni-Perception Loss compares the policy under full and unimodally masked inputs, applying a KL penalty only to modality-specific evidence tokens to suppress cross-modal hallucination. We further introduce MEP-Bench, a diagnostic benchmark that quantifies utilization and faithfulness. Experiments show that OPPO achieves state-of-the-art performance on MER-UniBench and MME-Emotion, while substantially improving utilization and faithfulness scores on MEP-Bench, highlighting the importance of sufficient and faithful omni perception for multimodal emotion reasoning.
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Emotion Recognition | IEMOCAP | -- | 151 | |
| Basic Emotion Recognition | MER 2023 | Hit Rate87.73 | 33 | |
| Multimodal Emotion Reasoning | MER-UniBench Aggregate | Mean Score81.05 | 14 | |
| Sentiment Analysis | MOSI | Sentiment Score86.5 | 14 | |
| Sentiment Analysis | SIMS V2 | Score88.26 | 14 | |
| Basic Emotion Recognition | MER 24 | Overall Score90.34 | 14 | |
| Basic Emotion Recognition | MELD | Accuracy64.06 | 14 | |
| Fine-grained Emotion Recognition | OV-MERD+ | Score67.16 | 14 | |
| Sentiment Analysis | SIMS | Score86.22 | 14 | |
| Multimodal Emotion Reasoning | MME Emotion | ER-Lab54.6 | 11 |