Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Plan, Watch, Recover: A Benchmark and Architectures for Proactive Procedural Assistance

About

We envision a proactive multi-modal assistant system which gives users real-time step-by-step guidance on a procedural task, autonomously deciding \textit{when} to interrupt, and \textit{how} to coach. However, progress is limited by the absence of large-scale, cross-domain benchmarks that reflect realistic conditions, particularly the common case in which users deviate from the expected step sequence. We address this gap with four contributions: \textbf{(1)}~we release \textbf{EgoProactive}, a large-scale wearable-egocentric dataset for proactive procedural assistance with explicit Out-of-Plan (OOP) annotations and recovery steps; \textbf{(2)}~we augment five established benchmarks (Ego4D, EPIC-KITCHENS, EgoExo4D, HoloAssist, HowTo100M) into \textbf{Pro\textsuperscript{2}Bench} under a unified proactive-guidance schema; \textbf{(3)}~we propose a \textbf{decoupled planner--interaction architecture} specialized for procedural state, visual cues, and recovery injection; \textbf{(4)}~we introduce a post-training recipe that transfers across model families, validated by cross-backbone replication on Llama~4 and Qwen-3.6-VL. In extensive experiments, our trained Llama-4 system substantially improves objective intervention quality over strong proprietary baselines (Claude Opus~4.6, Gemini~3.1~Pro, GPT~5.2) and open-weight baselines (Qwen3~VL~235B) baselines across all six datasets. Oracle-plan experiments further show that, when plan quality is controlled, the trained duplex model produces high-quality guidance and large gains on Out-of-Plan recovery.

Kaustav Kundu, Ritvik Shrivastava, Maxim Arap, Nanshu Wang, Xianhui Zhu, Quintin Fettes, Gautam Tiwari, Parth Suresh, Th\'eo Moutakanni, Alejandro Castillejo Munoz, Allen Bolourchi, Pascale Fung, Pinar Donmez, Babak Damavandi, Anuj Kumar, Seungwhan Moon• 2026

Related benchmarks

TaskDatasetResultRank
Proactive AssistancePro2Bench EE4D
G-Mean F189
10
Proactive AssistancePro2Bench EK
G-Mean F10.9
10
Proactive AssistancePro2Bench HowTo
G-Mean F188
10
Proactive AssistancePro2Bench Ego4D
G-Mean F186
10
Proactive AssistancePro2Bench HoloAssist
G-Mean F187
10
Proactive AssistancePro2Bench EgoProactive
G-Mean F168
10
Proactive AssistancePro2Bench Average
G-Mean F183
10
Subjective Guidance QualityEE4D
Subjective Guidance Quality (1-5)3.33
9
Subjective Guidance QualityEK
Subjective Guidance Quality3.67
9
Subjective Guidance QualityHowTo
Subjective Quality Score3.93
9
Showing 10 of 13 rows

Other info

Follow for update