SAC$^2$-Net: Semantic Anchoring and Complementary-Consensus Fusion for Multimodal Micro-Expression Recognition
About
Micro-expression recognition (MER) is challenging due to subtle facial movements, limited data, and the ambiguous relationship between Action Units (AUs) and emotion categories. Optical flow and motion magnification are two widely used representations for making subtle facial dynamics observable. However, many existing methods treat them as separate cues or fuse them without explicitly modeling their dual complementarity. Optical flow encodes displacement-level muscle motion, whereas motion magnification reveals appearance-level changes in facial texture and context. When both modalities are informative, their combination provides a more complete characterization of subtle facial dynamics; when one modality degrades, the other may still preserve discriminative evidence for compensation. This dual complementarity provides richer facial representations, but also introduces two key challenges for multimodal fusion: cross-modal heterogeneity and spatially varying modality reliability. To address these challenges, we propose SAC$^2$-Net, a Semantic Anchoring and Complementary-Consensus Network that first aligns heterogeneous visual representations with semantic anchors and then performs reliability-aware complementary fusion. Specifically, Semantic Anchoring Soft Alignment (SASA) converts activated AUs into textual prompts and uses hierarchical AU-aware soft labels to align motion-magnified and optical-flow representations while preserving semantic proximity among anatomically related samples. Based on the aligned representations, Complementary-Consensus Fusion (CCF) exchanges complementary motion and appearance cues, adaptively enhances unreliable local responses with trustworthy cross-modal evidence, and further encourages a shared spatial focus through consensus refinement.
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Micro-expression recognition | CASME II 3-class | Unweighted F1 Score97.33 | 86 | |
| Micro-expression recognition | SAMM 3-class | Unweighted F1 Score (UF1)89.17 | 60 | |
| Micro-expression recognition | CASME II 5-class | UF187.87 | 58 | |
| Micro-expression recognition | SAMM 5-class | UF1 Score82.09 | 55 | |
| Micro-expression recognition | SMIC-HS 3-class | UF184.22 | 23 | |
| 7-class micro-expression recognition | CAS(ME)3 | UAR58.28 | 23 | |
| 4-class micro-expression recognition | CAS(ME)3 | UF172.33 | 21 | |
| Micro-expression recognition | MEGC2019-CD 3-class setting | UF189.73 | 13 | |
| Micro-expression recognition | CASME II to SMIC 3-classes (test) | Accuracy56.46 | 12 | |
| Micro-expression recognition | SAMM to SMIC 3-classes (test) | Accuracy52.17 | 12 |