Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

End-to-End Facial Expression Detection in Long Videos

About

Facial expression detection requires spotting when expressions occur and recognizing which emotional category they belong to. Despite their close relationships, existing approaches typically address these tasks separately, limiting performance and robustness in real-world settings. In this work, we propose FEDN, a Facial Expression Detection Network, which unifies spotting and recognition into a single detection task performed fully end-to-end. FEDN introduces two temporal attention modules, segment-level attention to capture fine-grained local dynamics and sliding window attention to capture the broader temporal context. Their output is combined in a multi-scale temporal feature pyramid, which enables spotting of expressions with varying duration. This unified framework enables joint optimization and shared representation learning across tasks. FEDN outperforms strong baselines in both spotting and detection on three public benchmarks, demonstrating the effectiveness of unifying spotting and recognition across multiple temporal scales. Additionally, we uncover a previously unreported discrepancy between expert-annotated and self-reported emotion labels, highlighting a key challenge in expression benchmarking and motivating the development of more nuanced annotation protocols.

Yini Fang, Alec F. Diallo, Frederic Jumelle, Bertram Shi• 2025

Related benchmarks

TaskDatasetResultRank
Facial Expression SpottingCAS(ME)2--
16
Facial Expression RecognitionCAS(ME)² Annotated labels
F1 Score75.2
7
Facial Expression RecognitionCAS(ME)² Self-reported labels
F1 Score74.6
7
Facial Expression SpottingCAS(ME)² Annotated labels
F1 Score51.1
7
Facial Expression SpottingCAS(ME)² Self-reported labels
F1 Score52.4
7
Facial Expression SpottingCAS(ME)² Annotated labels
F1 Score35
7
SpottingSAMMLV
F1 Score44.7
7
Showing 7 of 7 rows

Other info

Follow for update