Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

GeoFace: Consistent Multi-View Face Generation with Geometry-Constrained Diffusion

About

We present GeoFace, a geometry-constrained multi-view diffusion framework for consistent face generation from a single input. % While recent multi-view diffusion models achieve photorealistic synthesis at the per-view level, they lack an explicit mechanism to enforce a shared 3D structure across views, often leading to inconsistent geometry across viewpoints. To address this, GeoFace proposes a unified dual-stream framework for joint generation of multi-view RGB images and 3D face geometry, where the appearance and geometry streams interact through shared attention layers. To encourage the two streams to mutually constrain each other, we introduce a geometry-guided attention alignment loss that supervises the cross-attention between appearance and geometry tokens with 3D-consistent correspondences, enabling the appearance stream to correctly reference pose-invariant geometric cues for robust alignment across viewpoints. Geometry is represented as a canonical UV position map, derived from a FLAME mesh fitted to multi-view observations, serving as a view-invariant shared constraint across all generated views. Experiments on RenderMe-360 and NeRSemble demonstrate that GeoFace consistently outperforms existing methods in both visual quality and cross-view geometric consistency, facilitating more efficient 3D reconstruction.

Yeji Choi, Jinhyeok Choi, Jaewon Min, Minkyung Kwon, Jin Hyeon Kim, Seungryong Kim• 2026

Related benchmarks

TaskDatasetResultRank
Novel View SynthesisNeRSemble v2 (test)
LPIPS0.2055
13
Novel View SynthesisRenderMe360 Frontal View
CSIM0.8084
12
Novel View SynthesisRenderMe-360 profile views (±45° to ±90°)
PSNR15.3
6
Novel View SynthesisNersemble frontal-to-mid views v2 (test)
PSNR21.33
6
Showing 4 of 4 rows

Other info

Follow for update