Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

THFM: A Unified Video Foundation Model for 4D Human Perception and Beyond

About

We present THFM, a unified video foundation model for human-centric perception that jointly addresses dense tasks (depth, normals, segmentation, dense pose) and sparse tasks (2d/3d keypoint estimation) within a single architecture. THFM is derived from a pretrained text-to-video diffusion model, repurposed as a single-forward-pass perception model and augmented with learnable tokens for sparse predictions. Modulated by the text prompt, our single unified model is capable of performing various perception tasks. Crucially, our model is on-par or surpassing state-of-the-art specialized models on a variety of benchmarks despite being trained exclusively on synthetic data (i.e.~without training on real-world or benchmark specific data). We further highlight intriguing emergent properties of our model, which we attribute to the underlying diffusion-based video representation. For example, our model trained on videos with a single human in the scene generalizes to multiple humans and other object classes such as anthropomorphic characters and animals -- a capability that hasn't been demonstrated in the past.

Letian Wang, Andrei Zanfir, Eduard Gabriel Bazavan, Misha Andriluka, Cristian Sminchisescu• 2026

Related benchmarks

TaskDatasetResultRank
Surface Normal EstimationHi4D
MAE11.01
32
3D Human Pose and Shape EstimationRICH 24 joints (test)
PA-MPJPE34.7
27
3D Human Pose and Shape EstimationHuman3.6M 14 joints (test)
PA-MPJPE27.7
20
Depth EstimationGoliath Face 12 cameras, 16 frames, 4 subjects 47
RMSE0.021
15
Depth EstimationGoliath FullBody 12 cameras, 16 frames, 4 subjects 47
RMSE0.037
15
Depth EstimationGoliath UpperBody 47 (12 cameras, 16 frames, 4 subjects)
RMSE0.028
15
3D keypoints reconstructionEMDB-1 24 joints (test)
PA-MPJPE35.2
12
Soft foreground segmentationPhotoMatte85
MSE9.00e-4
7
Soft foreground segmentationVideoMatte (static)
MAD5.24
6
Soft foreground segmentationPPM-100
MSE0.0044
6
Showing 10 of 10 rows

Other info

Follow for update