Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Helping Figures Tell their Story! Paper-Grounded Video Generation Explaining Complex Scientific Figures

About

Scientific figures compress complex pipelines into a single canvas, yet understanding them requires paper-grounded, step-by-step narration aligned with visual highlights a capability missing from current video generation systems and benchmarks. To address this, we introduce paper-grounded figure-to-video generation: generating narrated, region-grounded walkthrough videos from a figure and its paper. We propose MINARD (Multimodal Interpretation of Narrated Architecture via Region Decomposition), a pipeline that generates paper-grounded narrations and sequentially grounds them to figure regions. We also release FigTalk, a benchmark with new sequential and component-level grounding metrics derived. On FigTalk, MINARD generates humanlike, paper-faithful narrations and outperforms narration-conditioned figure spatial grounding compared to existing approaches in both automatic and human evaluation

Ishani Mondal, Javad Baghirov, Jordan Boyd-Graber• 2026

Related benchmarks

TaskDatasetResultRank
GroundingFigTalk-Gold All (reference set)
Macro F174.3
9
GroundingFigTalk-Gold Easy (reference set)
Macro F177.8
9
GroundingFigTalk-Gold Medium (reference set)
Macro F170.4
9
GroundingFigTalk-Gold Hard (reference set)
Macro F174.2
9
Narration Quality EvaluationFigTalk-Gold (D1)
Order Match76
9
Grounding EvaluationFigTalk Extended
Inputs CF Score84
8
Showing 6 of 6 rows

Other info

Follow for update