Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

GaussianEmoTalker: Real-Time Emotional Talking Head Synthesis with Audio-Driven and Blendshape-Based 3D Gaussian Splatting

About

Audio-driven talking head synthesis has achieved impressive progress in lip synchronization and visual quality, yet generating expressive emotional avatars with controllable intensity remains challenging, especially under real-time constraints. In this paper, we present GaussianEmoTalker, an audio-driven framework for real-time emotional talking head synthesis based on 3D Gaussian Splatting. Instead of directly predicting the final emotional avatar from speech, we formulate emotional animation as a neutral-to-emotional residual deformation problem. GaussianEmoTalker first constructs an identity-specific neutral talking space with GaussianBlendshapes, which provides high-fidelity Gaussian attributes and phoneme-synchronized neutral motion. It then predicts an emotion-conditioned residual deformation by combining mesh displacement cues, audio features, emotion categories, and intensity encodings. To fuse these heterogeneous signals, we introduce a spatial-audio-emotion attention module that estimates the offsets of Gaussian attributes for expressive and temporally stable rendering. Extensive experiments demonstrate that GaussianEmoTalker achieves competitive video quality, accurate lip synchronization, controllable emotional expression, and real-time rendering compared with recent emotional talking head methods. Our project page is available at https://njust-yang.github.io/GaussianEmoTalker.github.io/

Haijie Yang, Zhenyu Zhang, Yixuan Dong, Jianjun Qian, Jian Yang• 2026

Related benchmarks

TaskDatasetResultRank
Talking Head GenerationMEAD self-driven
FID26.36
10
Audio-Video SynchronizationCross-driven (Audio I)
Sync Error4.589
10
Audio-Video SynchronizationCross-driven Audio II
Sync4.478
10
Audio-Video SynchronizationCross-driven (Audio III)
Sync Error4.712
10
Audio-Video SynchronizationCross-driven (Audio IV)
Sync Error4.823
10
Audio-visual talking head generationUser Study Cross-driven setting
Emotional Accuracy40
6
Showing 6 of 6 rows

Other info

Follow for update