Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

EFlow: Learning Evidence Flow for Long-Video Reasoning with Adaptive Reflection

About

Long-video reasoning is fundamentally constrained by how models acquire and utilize visual evidence. Existing tool-augmented video frameworks often interleave temporal grounding and answer reasoning within a single trajectory, causing early semantic hypotheses to bias evidence localization. We term this failure mode premature semantic commitment, where biased grounding retrieves incomplete evidence and incomplete evidence further reinforces incorrect reasoning. To address this issue, we propose EFlow, an evidence-first video reasoning framework built upon Qwen3-VL. EFlow explicitly separates temporal grounding and logical reasoning through CoT for Temporal Grounding and CoT for Reasoning, enabling the model to retrieve relevant evidence before answer inference. In addition, EFlow introduces a confidence-aware reflection mechanism that re-evaluates the full video when retrieved evidence is potentially insufficient. We further construct dedicated trajectory datasets and train EFlow through supervised fine-tuning, reinforcement learning, and reinforcement fine-tuning. Extensive experiments across five video understanding benchmarks demonstrate that EFlow consistently improves long-video reasoning performance.

Wenhao Zhang, Kuanwei Lin, Xuyi Yang, Wei Gao, Ge Li• 2026

Related benchmarks

TaskDatasetResultRank
Video ReasoningVSI-Bench
Accuracy59.1
101
Video ReasoningVideo-MMMU
Accuracy65.5
83
Video Question AnsweringVideo-MME Long
Accuracy60.1
71
Video Question AnsweringVideoMME Overall
Accuracy69.1
40
Video Question AnsweringLongVideoBench (standard)
Accuracy60.3
19
Video Question AnsweringLVBench Avg
Average Accuracy52.5
11
Temporal Video GroundingNextGQA (Overall)
Overall Accuracy80
5
Showing 7 of 7 rows

Other info

Follow for update