Cambrian-P: Pose-Grounded Video Understanding

About

Camera pose matters. The position and orientation of each viewpoint define a shared spatial coordinate frame that relates observations across video frames. Yet this signal is largely absent from multimodal LLMs (MLLMs) for video understanding, which process frames as isolated 2D snapshots, instead of the persistent scene humans perceive. We revisit pose as a lightweight supervisory signal and introduce Cambrian-P, a video MLLM augmented with per-frame learnable camera tokens and a pose regression head. With a carefully designed sampling scheme, the model achieves substantial gains of 4.5-6.5% on spatial reasoning benchmarks such as VSI-Bench, generalizes across eight additional spatial and general video QA benchmarks, and, as a byproduct, achieves state of the art streaming pose estimation on ScanNet. Surprisingly, training on pseudo-annotated poses from in-the-wild video further improves general video QA benchmarks, showing pose helps beyond spatial reasoning. Together, these results position camera pose as a fundamental signal for video models that reason about the physical world.

Jihan Yang, Zifan Zhao, Xichen Pan, Shusheng Yang, Junyi Zhang, Bingyi Kang, Hu Xu, Shang-Wen Li, Saining Xie• 2026

Related benchmarks

Task	Dataset	Result
Spatial Reasoning	VSI-Bench	R.Dr.89.5	370
Camera pose estimation	TUM-dynamic	ATE0.046	205
Camera pose estimation	ScanNet	RPE (t)0.023	133
Spatial Reasoning	ReVSI	Average Score52	50
Camera pose estimation	Sintel dataset	ATE0.239	35
Temporal spatial reasoning	VSTI-Bench (test)	Average Score68.9	20
Camera pose estimation	ScanNet (test)	Per-sequence Time (s)2.16	6

Showing 7 of 7 rows

Other info

Follow for update

@wizwand_team Discord