Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

SplatReasoner: Enhancing Embodied Reasoning and Grounding by Novel View Synthesis

About

Vision-Language Models (VLMs) have demonstrated strong reasoning capabilities over images and videos, yet their application to embodied scene understanding often constrained by the fixed viewpoints stored in episodic RGB-D memories. These observations may fail to capture query-relevant evidence due to occlusions, object truncation, restricted fields of view, or suboptimal view composition. We present SplatReasoner, a framework that introduces novel view synthesis into the VLM reasoning process by leveraging 3D Gaussian Splatting (3DGS). Given a user query about a 3D scene, SplatReasoner retrieves relevant observations and synthesizes query-conditioned viewpoints that reveal the visual evidence needed to answer the query and ground the referred entities in 3D. Experiments show that query-conditioned novel view synthesis improves both embodied reasoning and 3D grounding over fixed-view memory and language-embedded 3DGS baselines.

Kim Yu-Ji, Dahye Lee, Kim Jun-Seong, Nam Hyeon-Woo, GeonU Kim, Yongjin Kwon, Yu-Chiang Frank Wang, Jaesung Choe, Tae-Hyun Oh• 2026

Related benchmarks

TaskDatasetResultRank
Embodied Question AnsweringOpenEQA EM-EQA
LLM-Match57.8
8
3D Referring SegmentationScanNet curated (test)
3D mIoU12.87
5
Showing 2 of 2 rows

Other info

Follow for update