Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Conan: Progressive Learning to Reason Like a Detective over Multi-Scale Visual Evidence

About

Video reasoning, which requires multi-step deduction across frames, remains a major challenge for multimodal large language models (MLLMs). While reinforcement learning (RL)-based methods enhance reasoning capabilities, they often rely on text-only chains that yield ungrounded or hallucinated conclusions. Conversely, frame-retrieval approaches introduce visual grounding, yet still struggle with inaccurate evidence localization. To address these limitations, we present Conan, a framework for evidence-grounded multi-step video reasoning. Conan identifies context and evidence frames, reasons over cross-frame clues, and adaptively decides when to conclude or explore further. To achieve this, we 1) construct Conan-91K, a large-scale dataset of automatically generated reasoning traces that include frame identification, evidence reasoning, and action decision, and 2) design a multi-stage progressive cold-start strategy combined with an Identification-Reasoning-Action (AIR) RLVR training framework to progressively incentivize multi-step visual reasoning. Extensive experiments on six multi-step reasoning benchmarks demonstrate that Conan surpasses the baseline Qwen2.5-VL-7B-Instruct by an average of over 10% in accuracy, achieving state-of-the-art performance. Furthermore, Conan generalizes effectively to long video understanding tasks, validating its strong scalability and robustness.

Kun Ouyang, Yuanxin Liu, Linli Yao, Yishuo Cai, Hao Zhou, Jie Zhou, Fandong Meng, Xu Sun• 2025

Related benchmarks

TaskDatasetResultRank
Long Video UnderstandingMLVU--
265
Video Question AnsweringVideoMME--
254
Video Question AnsweringLongVideoBench
Accuracy56.6
224
Video UnderstandingMLVU
Accuracy59.2
147
Video UnderstandingLongVideoBench
Accuracy54.5
128
Video UnderstandingLVBench
Overall Accuracy38.2
95
Video UnderstandingMMVU
Accuracy64
91
Long Video UnderstandingVideo-MME Overall
Accuracy60.5
81
Video Question AnsweringMLVU
M-Avg Score63.4
80
Temporal GroundingCharades-STA (test)--
68
Showing 10 of 16 rows

Other info

Follow for update