EMOPortraits: Emotion-enhanced Multimodal One-shot Head Avatars

About

Head avatars animated by visual signals have gained popularity, particularly in cross-driving synthesis where the driver differs from the animated character, a challenging but highly practical approach. The recently presented MegaPortraits model has demonstrated state-of-the-art results in this domain. We conduct a deep examination and evaluation of this model, with a particular focus on its latent space for facial expression descriptors, and uncover several limitations with its ability to express intense face motions. To address these limitations, we propose substantial changes in both training pipeline and model architecture, to introduce our EMOPortraits model, where we: Enhance the model's capability to faithfully support intense, asymmetric face expressions, setting a new state-of-the-art result in the emotion transfer task, surpassing previous methods in both metrics and quality. Incorporate speech-driven mode to our model, achieving top-tier performance in audio-driven facial animation, making it possible to drive source identity through diverse modalities, including visual signal, audio, or a blend of both. We propose a novel multi-view video dataset featuring a wide range of intense and asymmetric facial expressions, filling the gap with absence of such data in existing datasets.

Nikita Drobyshev, Antoni Bigata Casademunt, Konstantinos Vougioukas, Zoe Landgraf, Stavros Petridis, Maja Pantic• 2024

Related benchmarks

Task	Dataset	Result
Portrait Animation (Self-reenactment)	VFHQ (test)	FVD444.5	23
Controllable Image Generation and Editing	CelebA-HQ (test)	Accuracy68.9	20
Human Image Controllability and Editing	AffectHuman-43K (test)	Accuracy69.8	20
Facial Image Editing	AffectNet	Accuracy63.8	20
Self-reenactment portrait animation	MEAD 59 (test)	CSIM0.5959	18
Portrait Animation (Cross-reenactment)	FFHQ source + VFHQ driving (test)	CSIM0.3483	18
Talking head video generation	HDTF	FID24.77	14
Video-driven Talking Head Generation (Self-Reenactment)	HDTF	FID27.71	12
Talking head video generation	Talkinghead1kh	FID41.08	8
Cross-Reenactment	FFHQ 512x512	FID105.1	6

Showing 10 of 20 rows

Other info

Follow for update

@wizwand_team Discord