Action-Conditioned 3D Human Motion Synthesis with Transformer VAE

About

We tackle the problem of action-conditioned generation of realistic and diverse human motion sequences. In contrast to methods that complete, or extend, motion sequences, this task does not require an initial pose or sequence. Here we learn an action-aware latent representation for human motions by training a generative variational autoencoder (VAE). By sampling from this latent space and querying a certain duration through a series of positional encodings, we synthesize variable-length motion sequences conditioned on a categorical action. Specifically, we design a Transformer-based architecture, ACTOR, for encoding and decoding a sequence of parametric SMPL human body models estimated from action recognition datasets. We evaluate our approach on the NTU RGB+D, HumanAct12 and UESTC datasets and show improvements over the state of the art. Furthermore, we present two use cases: improving action recognition through adding our synthesized data to training, and motion denoising. Code and models are available on our project page.

Mathis Petrovich, Michael J. Black, G\"ul Varol• 2021

Related benchmarks

Task	Dataset	Result
3D Human Motion Generation	HumanAct12	FID0.12	36
3D Human Motion Generation	UESTC	Accuracy91.82	14
Action-conditional motion synthesis	UESTC (train test)	FID (Train Set)20.5	13
Action-conditioned Motion Generation	CMU MOCAP	Accuracy15.21	10
Unconditional Motion Generation	HumanML3D (test)	FID14.14	10
3D Human Pose Forecasting	3DPW	MPJPE69.4	9
Human Action Recognition	NTU RGB+D 120 Few-Shot	Median Top-1 Acc73.6	9
Bi-ventricular surface reconstruction	ACDC + M&Ms + M&Ms-2 (test)	ASSD (LV)2.1	7
Unconditional human motion synthesis	HumanAct12	FID48.8	7
Unconstrained Synthesis	HumanAct12 (unconstrained)	FID48.8	7

Showing 10 of 20 rows

Other info

Follow for update

@wizwand_team Discord