Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Claim-Level Rubric Rewards for Video Caption Reinforcement Learning

About

In this paper, we introduce Claim-Level Rubric Rewards (CuRe), a structured reward framework designed to address the reward-design bottleneck in reinforcement learning for dense video captioning. Existing reward designs generally fall into two categories: holistic response-level judgment across heterogeneous criteria, or alignment-based evaluation against reference captions. However, both paradigms suffer from fundamental limitations. Holistic rewards struggle to ensure factual accuracy and are prone to stylistic reward hacking, while reference-based rewards overly rely on rigid textual alignment, failing to preserve the completeness and diversity inherent to open-ended generation tasks. To address these challenges, CuRe reformulates reward modeling as fine-grained claim-level verification. Specifically, CuRe decomposes captions into category-aware atomic claims through a structured rubric, converting holistic evaluation into simpler and more reliable claim-level verification.

Mingqi Gao, Hongyuan Dong, Yifei Chen, Zhisheng Zhong, Zheng Ruan, Wenjin Hou, Yu Chen, Han Hu, Yansong Tang• 2026

Related benchmarks

TaskDatasetResultRank
Video Question AnsweringPerceptionTest
Accuracy63.56
39
Video Question AnsweringVideoMME
Short VQA Accuracy73.67
32
Video CaptioningDREAM 1k
Precision48.21
21
Video CaptioningVCapsBench
AR67.77
17
Video Question AnsweringMotionBench
ALL47.86
11
Video CaptioningEventHallusion Description
Entire Hallucination Score45.87
11
Video Question AnsweringMVBench
Overall Score54.43
10
Video Question AnsweringTOMATO
Score22.91
8
Showing 8 of 8 rows

Other info

Follow for update