Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

MLCR: Multi-Level Cue Refinement for Long-Term Multimodal Action Quality Assessment

About

Long-term multimodal action quality assessment (AQA) evaluates action execution in several-minute audiovisual sequences by mining discriminative quality cues for score prediction. Existing multimodal methods usually model entire sequences with a single temporal encoder and fuse modality features by direct alignment or concatenation, causing key cues to be obscured by global trends, weakened by modal redundancy, and distorted during one-shot score mapping. To address this issue, we reformulate long-term multimodal AQA as a quality cue organization problem and propose MLCR, a multi-level cue refinement framework. MLCR organizes quality evidence at three levels: intra-modal representation, cross-modal interaction, and stage-wise aggregation. Specifically, the intra-modal decoupling encoder (IMDE) preserves modality identity while refining global temporal context and local frequency details. The cross-modal dynamic complementarity-aware retrieval (CMDCR) module retrieves incremental evidence conditioned on the evolving fused state and suppresses redundant responses. The stage-wise multimodal integration (SMI) block progressively accumulates intra-modal and cross-modal cues to refine the fused representation. Experiments on the Rhythmic Gymnastics and Fis-V datasets show that MLCR achieves the best or second-best performance in both Spearman correlation and prediction error, demonstrating its effectiveness and robustness.

Qiqi Li, Pengfei Wang, Hongyu Chen, Nenggan Zheng• 2026

Related benchmarks

TaskDatasetResultRank
Action Quality AssessmentFis-V
TES Spearman Correlation0.744
22
Action Quality AssessmentRhythmic Gymnastics (test)
Spearman Correlation (Ball)0.808
16
Showing 2 of 2 rows

Other info

Follow for update