Our new X account is live! Follow @wizwand_team for updates
WorkDL logo mark

Skip-Plan: Procedure Planning in Instructional Videos via Condensed Action Space Learning

About

In this paper, we propose Skip-Plan, a condensed action space learning method for procedure planning in instructional videos. Current procedure planning methods all stick to the state-action pair prediction at every timestep and generate actions adjacently. Although it coincides with human intuition, such a methodology consistently struggles with high-dimensional state supervision and error accumulation on action sequences. In this work, we abstract the procedure planning problem as a mathematical chain model. By skipping uncertain nodes and edges in action chains, we transfer long and complex sequence functions into short but reliable ones in two ways. First, we skip all the intermediate state supervision and only focus on action predictions. Second, we decompose relatively long chains into multiple short sub-chains by skipping unreliable intermediate actions. By this means, our model explores all sorts of reliable sub-relations within an action sequence in the condensed action space. Extensive experiments show Skip-Plan achieves state-of-the-art performance on the CrossTask and COIN benchmarks for procedure planning.

Zhiheng Li, Wenjia Geng, Muheng Li, Lei Chen, Yansong Tang, Jiwen Lu, Jie Zhou• 2023

Related benchmarks

TaskDatasetResultRank
Procedure PlanningCrossTask
Success Rate (SR)28.85
35
Procedure PlanningCOIN T=3 (test)
SR0.2365
21
Procedure PlanningCrossTask T=5
Success Rate8.55
15
Procedure PlanningCOIN T=4 (test)
SR16.04
13
Procedure PlanningCrossTask long horizons T=6
Success Rate (SR)5.12
10
Procedure PlanningCOIN T=5 (test)
SR9.9
8
Showing 6 of 6 rows

Other info

Follow for update