SkillSpotter: Pose-Aware Multi-View Skilled Action Detection and Grading in Ego-Exo Videos
About
To enable personalized, real-time coaching using Augmented Reality glasses or fixed camera setups in domains such as sports, cooking, or music, a system must understand not just what a person does, but how well they execute an activity. In an ego-exo video setting, this requires simultaneously detecting individual skilled actions and classifying each as correct or needing improvement, which Ego-Exo4D's proficiency demonstration benchmark formalized. We first adapt seven state-of-the-art temporal action detection architectures to this task, extend the evaluation protocol to disentangle detection from grading, and show that existing methods grade near-randomly. We then introduce SkillSpotter, a pose-aware multi-view architecture that jointly detects and grades skilled actions through three task-specific modules: (1) adaptive temporal suppression to handle the varying density of skilled actions across diverse activities, (2) gated 3D body pose fusion to leverage body kinematics as a complementary signal to visual features, and (3) bidirectional cross-view attention to combine ego and exo views effectively. SkillSpotter improves class-specific mAP from 12.40 to 21.82 (+76%) and balanced accuracy from 55.99% to 60.40% over the best baseline. SkillSpotter's modules transfer to other temporal action detection models with consistent gains, and our method generalizes beyond Ego-Exo4D to HoloAssist. Code: https://github.com/eth-siplab/SkillSpotter
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Timestamp-level skill assessment | Ego-Exo4D Ego view 1.0 (demonstration proficiency) | mAPS21.82 | 12 | |
| Timestamp-level skill assessment | Ego-Exo4D Exos view 1.0 (demonstration proficiency) | mAPS21.12 | 12 | |
| Timestamp-level skill assessment | Ego-Exo4D Ego+Exos view 1.0 (demonstration proficiency) | mAPS21.34 | 12 | |
| Mistake detection | HoloAssist | F153.3 | 4 |