Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

MotionAtlas: Detailed Region Captioning for Motion-Centric Videos

About

We propose MotionAtlas, a system for detailed captioning of motion-centric videos, comprising (1) a dedicated human-annotated benchmark, (2) a scalable, high-quality pipeline to construct training samples, and (3) a family of powerful Video-MLLMs. Unlike conventional global motion captioning datasets, we focus on region-aware motion captioning: given a video and a spatiotemporal mask, the model generates precise descriptions of motion within the target region, thereby alleviating visual clutter and motion entanglement and enabling reliable, quantifiable evaluation. Concretely, we first build MotionAtlas-Bench, a comprehensive benchmark comprising 2,073 multiple-choice questions, meticulously annotated for a curated set of high-quality, motion-centric videos, to evaluate fine-grained motion understanding of the objects in question. Second, we design a rigorous and scalable data pipeline that leverages self-bootstrap refinement to suppress fine-grained hallucinations, yielding 159k high-quality motion captioning data. Third, we design a tailored training data composition strategy, which achieves consistent and substantial performance gains across diverse baseline Video-MLLMs, including Molmo2 and Qwen3-VL. For instance, MotionAtlas-4B surpasses Qwen3-VL-4B by an average of 5.2 percentage points across general motion benchmarks. The benchmark, dataset, and code have been released.

Weisong Liu, Haochen Wang, Kuan Gao, Yuhao Wang, Yikang Zhou, Zhongwei Ren, Jacky Mai, Anna Wang, Yanwei Li, Jason Li, Zhaoxiang Zhang• 2026

Related benchmarks

TaskDatasetResultRank
Temporal Video UnderstandingTempCompass
Accuracy75.7
160
Motion UnderstandingMotionBench
Accuracy62.6
35
General Video UnderstandingTVBench
Accuracy54.8
17
Single-Frame GroundingMotionAtlas-Bench
Overall Accuracy31.6
14
Full-Sequence GroundingMotionAtlas-Bench
Overall Accuracy34.1
14
Action-Attribute RecognitionFAVOR-Bench
Accuracy58.7
11
Action-Attribute RecognitionTOMATO
Accuracy36.8
11
Video UnderstandingDREAM 1k
F1-Score39.6
11
Showing 8 of 8 rows

Other info

Follow for update