Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

SER: Learning to Ground Video Reasoning with Semantic Evidence Rewards

About

Video MLLMs often struggle with fine-grained spatio-temporal reasoning, sometimes generating correct answers based on irrelevant frames or objects. Although outputting spatio-temporal evidence during reasoning is a promising direction, existing RL frameworks typically rely on geometry-only (IoU) rewards, which can be sensitive to boundary perturbations and overlook semantic alignment. To address this, we propose Semantic Evidence Reward (SER), which reformulates spatio-temporal evidence grounding as a constrained verification task. Instead of computing pixel-level overlap, SER uses a referee VLM as a local checker to evaluate model-generated evidence claims across two dimensions: relevance and localization quality, combined with a temporal penalty. This design reduces the reliance on dense box annotations and enables training directly on standard video QA data. On the V-STAR benchmark, SER achieves 49.6% mLGM, improving by 3.0 points over the strong evidence-grounded baseline Open-o3-Video, demonstrating its potential in enhancing both answer accuracy and evidence grounding.

Sheng Xia, Zhengqin Lai, Tianxiang Jiang, Kanghui Tian, Shoujun Zhou, Bin Li, Yi Wang• 2026

Related benchmarks

TaskDatasetResultRank
Temporal GroundingActivityNet
Recall@0.358.1
111
Video ReasoningLongVideoReason
Accuracy71.7
61
Multi-modal Video EvaluationVideoMME
Score62.9
50
Video UnderstandingWorldSense
Score42.3
33
Spatio-Temporal ReasoningV-STAR (test)
What Accuracy61.6
27
Multimodal Video UnderstandingVideoMMMU
Overall Score51.4
25
Temporal Video GroundingTVGBench
mIoU26.5
20
Showing 7 of 7 rows

Other info

Follow for update