3D Gaussian Blendshapes for Head Avatar Animation
About
We introduce 3D Gaussian blendshapes for modeling photorealistic head avatars. Taking a monocular video as input, we learn a base head model of neutral expression, along with a group of expression blendshapes, each of which corresponds to a basis expression in classical parametric face models. Both the neutral model and expression blendshapes are represented as 3D Gaussians, which contain a few properties to depict the avatar appearance. The avatar model of an arbitrary expression can be effectively generated by combining the neutral model and expression blendshapes through linear blending of Gaussians with the expression coefficients. High-fidelity head avatar animations can be synthesized in real time using Gaussian splatting. Compared to state-of-the-art methods, our Gaussian blendshape representation better captures high-frequency details exhibited in input video, and achieves superior rendering performance.
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Self-Reenactment | HDTF | PSNR27.81 | 35 | |
| Self-Reenactment | INSTA | PSNR29.64 | 19 | |
| Audio-Video Synchronization | Cross-driven (Audio IV) | Sync Error4.149 | 10 | |
| Audio-Video Synchronization | Cross-driven Audio II | Sync4.459 | 10 | |
| Audio-Video Synchronization | Cross-driven (Audio III) | Sync Error4.648 | 10 | |
| Audio-Video Synchronization | Cross-driven (Audio I) | Sync Error4.762 | 10 | |
| Talking Head Generation | MEAD self-driven | FID32.33 | 10 | |
| Head Avatar Reconstruction | INSTA dataset (test) | PSNR (bala)33.21 | 8 | |
| Head Avatar Reconstruction | GaussianBlendShapes (test) | PSNR (Subject 1)33.14 | 8 | |
| Head Avatar Rendering | INSTA | Inverse MAE98 | 7 |