Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

GenZI: Zero-Shot 3D Human-Scene Interaction Generation

About

Can we synthesize 3D humans interacting with scenes without learning from any 3D human-scene interaction data? We propose GenZI, the first zero-shot approach to generating 3D human-scene interactions. Key to GenZI is our distillation of interaction priors from large vision-language models (VLMs), which have learned a rich semantic space of 2D human-scene compositions. Given a natural language description and a coarse point location of the desired interaction in a 3D scene, we first leverage VLMs to imagine plausible 2D human interactions inpainted into multiple rendered views of the scene. We then formulate a robust iterative optimization to synthesize the pose and shape of a 3D human model in the scene, guided by consistency with the 2D interaction hypotheses. In contrast to existing learning-based approaches, GenZI circumvents the conventional need for captured 3D interaction data, and allows for flexible control of the 3D interaction synthesis with easy-to-use text prompts. Extensive experiments show that our zero-shot approach has high flexibility and generality, making it applicable to diverse scene types, including both indoor and outdoor environments.

Lei Li, Angela Dai• 2023

Related benchmarks

TaskDatasetResultRank
HHOI generationHHOI
HH Penetration Ratio (x1000)13.91
12
3D Human-Scene InteractionPROX-s filtered
Semantic Clip Quality0.2521
4
General Human-scene InteractionSceneFun3D curated subset
SCS0.2542
3
text-guided dyadic HHOI generationCORE4D and authors' collected dataset (test)
Body Pose Fidelity1.35
3
Multi-human generationText-guided Multi-human HHOI Generation
CLIP Score0.2584
3
Functional Human-scene InteractionSceneFun3D curated subset
SCS25.01
3
Showing 6 of 6 rows

Other info

Code

Follow for update