Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

HOI-Diff: Text-Driven Synthesis of 3D Human-Object Interactions using Diffusion Models

About

We address the problem of generating realistic 3D human-object interactions (HOIs) driven by textual prompts. To this end, we take a modular design and decompose the complex task into simpler sub-tasks. We first develop a dual-branch diffusion model (HOI-DM) to generate both human and object motions conditioned on the input text, and encourage coherent motions by a cross-attention communication module between the human and object motion generation branches. We also develop an affordance prediction diffusion model (APDM) to predict the contacting area between the human and object during the interactions driven by the textual prompt. The APDM is independent of the results by the HOI-DM and thus can correct potential errors by the latter. Moreover, it stochastically generates the contacting points to diversify the generated motions. Finally, we incorporate the estimated contacting points into the classifier-guidance to achieve accurate and close contact between humans and objects. To train and evaluate our approach, we annotate BEHAVE dataset with text descriptions. Experimental results on BEHAVE and OMOMO demonstrate that our approach produces realistic HOIs with various interactions and different types of objects.

Xiaogang Peng, Yiming Xie, Zizhao Wu, Varun Jampani, Deqing Sun, Huaizu Jiang• 2023

Related benchmarks

TaskDatasetResultRank
Human-Object Interaction GenerationOMOMO (test)
FID0.245
24
Human-Object Interaction GenerationBEHAVE (test)
FID0.437
21
Human-Object Interaction SynthesisFullBodyManipulation (test)
FID11.25
19
3D Human-Object Interaction GenerationOMOMO (test)
FID0.245
9
3D Human-Object Interaction GenerationBEHAVE (test)
FID0.437
9
action-conditioned interaction generationInterAct (test)
FID3.566
9
Human-Object Interaction SynthesisOMOMO Seen Objects (test)
Contact Recall (Crec)62
7
Human-Object Interaction GenerationInterAct v1 (test)
R-Precision (Top 1)41.3
7
Human-Object Interaction SynthesisOMOMO (5 strictly unseen objects)
Recall (C)45
6
Text-to-Dynamic HOI Motion GenerationDedicated dynamic HOI (test)
Skating Quality Score63.3
5
Showing 10 of 15 rows

Other info

Follow for update