InstructScene: Instruction-Driven 3D Indoor Scene Synthesis with Semantic Graph Prior
About
Comprehending natural language instructions is a charming property for 3D indoor scene synthesis systems. Existing methods directly model object joint distributions and express object relations implicitly within a scene, thereby hindering the controllability of generation. We introduce InstructScene, a novel generative framework that integrates a semantic graph prior and a layout decoder to improve controllability and fidelity for 3D scene synthesis. The proposed semantic graph prior jointly learns scene appearances and layout distributions, exhibiting versatility across various downstream tasks in a zero-shot manner. To facilitate the benchmarking for text-driven 3D scene synthesis, we curate a high-quality dataset of scene-instruction pairs with large language and multimodal models. Extensive experimental results reveal that the proposed method surpasses existing state-of-the-art approaches by a large margin. Thorough ablation studies confirm the efficacy of crucial design components. Project page: https://chenguolin.github.io/projects/InstructScene.
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Text-to-scene generation | 3D-FRONT Livingroom (test) | FID110.4 | 15 | |
| Scene Rearrangement | 3D-FRONT Living room | FID106 | 13 | |
| Scene Rearrangement | 3D-FRONT Bedroom | FID105.3 | 13 | |
| Text-to-scene generation | 3D-FRONT Diningroom (test) | FID129.1 | 10 | |
| Text-to-scene generation | 3D-FRONT Bedroom (test) | FID114.9 | 10 | |
| Unconditional Scene Synthesis | 3D-FRONT Bedroom | FID125 | 10 | |
| Unconditional Scene Synthesis | 3D-FRONT Dining room | FID137.5 | 10 | |
| Unconditional Scene Synthesis | 3D-FRONT Living room | FID117.6 | 10 | |
| 3D Indoor Scene Synthesis | Conventional Manhattan environments standard (test) | OOB211.4 | 7 | |
| Scene Completion | 3D-FRONT Dining rooms | FID107.9 | 7 |