Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

STEPs: Self-Supervised Key Step Extraction and Localization from Unlabeled Procedural Videos

About

We address the problem of extracting key steps from unlabeled procedural videos, motivated by the potential of Augmented Reality (AR) headsets to revolutionize job training and performance. We decompose the problem into two steps: representation learning and key steps extraction. We propose a training objective, Bootstrapped Multi-Cue Contrastive (BMC2) loss to learn discriminative representations for various steps without any labels. Different from prior works, we develop techniques to train a light-weight temporal module which uses off-the-shelf features for self supervision. Our approach can seamlessly leverage information from multiple cues like optical flow, depth or gaze to learn discriminative features for key-steps, making it amenable for AR applications. We finally extract key steps via a tunable algorithm that clusters the representations and samples. We show significant improvements over prior works for the task of key step localization and phase classification. Qualitative results demonstrate that the extracted key steps are meaningful and succinctly represent various steps of the procedural tasks.

Anshul Shah, Benjamin Lundell, Harpreet Sawhney, Rama Chellappa• 2023

Related benchmarks

TaskDatasetResultRank
Procedure LearningProceL
Precision23.5
13
Procedure LearningCrossTask
Precision26.2
13
Procedure LearningEgoProceL EPIC-Tents
F1 Score42.2
11
Procedure LearningEgoProceL EGTEA-GAZE+
F1 Score30.8
11
Procedure LearningEgoProceL MECCANO
F1 Score36.4
11
Procedure LearningEgoProceL CMU-MMAC
F1 Score28.3
11
Procedure LearningEgoProceL PC Assembly
F1 Score24.9
11
Procedure LearningEgoProceL PC Disassembly
F1 Score25.9
11
Showing 8 of 8 rows

Other info

Follow for update