MotionAtlas: Detailed Region Captioning for Motion-Centric Videos
About
We propose MotionAtlas, a system for detailed captioning of motion-centric videos, comprising (1) a dedicated human-annotated benchmark, (2) a scalable, high-quality pipeline to construct training samples, and (3) a family of powerful Video-MLLMs. Unlike conventional global motion captioning datasets, we focus on region-aware motion captioning: given a video and a spatiotemporal mask, the model generates precise descriptions of motion within the target region, thereby alleviating visual clutter and motion entanglement and enabling reliable, quantifiable evaluation. Concretely, we first build MotionAtlas-Bench, a comprehensive benchmark comprising 2,073 multiple-choice questions, meticulously annotated for a curated set of high-quality, motion-centric videos, to evaluate fine-grained motion understanding of the objects in question. Second, we design a rigorous and scalable data pipeline that leverages self-bootstrap refinement to suppress fine-grained hallucinations, yielding 159k high-quality motion captioning data. Third, we design a tailored training data composition strategy, which achieves consistent and substantial performance gains across diverse baseline Video-MLLMs, including Molmo2 and Qwen3-VL. For instance, MotionAtlas-4B surpasses Qwen3-VL-4B by an average of 5.2 percentage points across general motion benchmarks. The benchmark, dataset, and code have been released.
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Temporal Video Understanding | TempCompass | Accuracy75.7 | 160 | |
| Motion Understanding | MotionBench | Accuracy62.6 | 35 | |
| General Video Understanding | TVBench | Accuracy54.8 | 17 | |
| Single-Frame Grounding | MotionAtlas-Bench | Overall Accuracy31.6 | 14 | |
| Full-Sequence Grounding | MotionAtlas-Bench | Overall Accuracy34.1 | 14 | |
| Action-Attribute Recognition | FAVOR-Bench | Accuracy58.7 | 11 | |
| Action-Attribute Recognition | TOMATO | Accuracy36.8 | 11 | |
| Video Understanding | DREAM 1k | F1-Score39.6 | 11 |