Claim-Level Rubric Rewards for Video Caption Reinforcement Learning
About
In this paper, we introduce Claim-Level Rubric Rewards (CuRe), a structured reward framework designed to address the reward-design bottleneck in reinforcement learning for dense video captioning. Existing reward designs generally fall into two categories: holistic response-level judgment across heterogeneous criteria, or alignment-based evaluation against reference captions. However, both paradigms suffer from fundamental limitations. Holistic rewards struggle to ensure factual accuracy and are prone to stylistic reward hacking, while reference-based rewards overly rely on rigid textual alignment, failing to preserve the completeness and diversity inherent to open-ended generation tasks. To address these challenges, CuRe reformulates reward modeling as fine-grained claim-level verification. Specifically, CuRe decomposes captions into category-aware atomic claims through a structured rubric, converting holistic evaluation into simpler and more reliable claim-level verification.
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Video Question Answering | PerceptionTest | Accuracy63.56 | 39 | |
| Video Question Answering | VideoMME | Short VQA Accuracy73.67 | 32 | |
| Video Captioning | DREAM 1k | Precision48.21 | 21 | |
| Video Captioning | VCapsBench | AR67.77 | 17 | |
| Video Question Answering | MotionBench | ALL47.86 | 11 | |
| Video Captioning | EventHallusion Description | Entire Hallucination Score45.87 | 11 | |
| Video Question Answering | MVBench | Overall Score54.43 | 10 | |
| Video Question Answering | TOMATO | Score22.91 | 8 |