Procedure Learning via Regularized Gromov-Wasserstein Optimal Transport
About
We study self-supervised procedure learning, which discovers key steps and their order from a set of unlabeled videos. Previous methods typically learn frame-to-frame correspondences between videos before determining key steps and their order. However, their performance often suffers from order variations, background/redundant frames, and repeated actions. To overcome these challenges, we propose a self-supervised framework, which utilizes a fused Gromov-Wasserstein optimal transport with a structural prior for frame-to-frame mapping. However, optimizing only for the above temporal alignment may lead to degenerate solutions, where all frames are mapped to a small cluster in the embedding space and thus every video is assigned to just one key step. To address that issue, we integrate a contrastive regularization, which maps different frames to various points, avoiding trivial solutions. Finally, extensive experiments on egocentric and third-person benchmarks demonstrate our superior performance over prior works, including OPEL which relies on a classical Kantorovich optimal transport with an optimality prior.
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Procedure Learning | ProceL | Precision42.2 | 13 | |
| Procedure Learning | CrossTask | Precision40.4 | 13 | |
| Procedure Learning | EgoProceL PC Assembly | F1 Score43.6 | 11 | |
| Procedure Learning | EgoProceL PC Disassembly | F1 Score45.9 | 11 | |
| Procedure Learning | EgoProceL MECCANO | F1 Score59.5 | 11 | |
| Procedure Learning | EgoProceL CMU-MMAC | F1 Score54.4 | 11 | |
| Procedure Learning | EgoProceL EGTEA-GAZE+ | F1 Score37.4 | 11 | |
| Procedure Learning | EgoProceL EPIC-Tents | F1 Score39.7 | 11 | |
| Action Segmentation | ProceL | F1 Score44.3 | 9 | |
| Action Segmentation | CrossTask | F1 Score40.4 | 9 |