Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

SATURN: Symbolic Spatial Reasoning for Multi-Perspective Grounding

About

Vision-Language Models (VLMs) remain unreliable when spatial reasoning requires composing relations whose meanings depend on frames of reference. Existing neuro-symbolic methods make reasoning more explicit, but often depend on brittle geometric procedures and hard decisions over noisy perception. We propose SATURN, a neuro-symbolic framework for perspective-aware compositional spatial reasoning. SATURN reconstructs an approximate 3D scene, derives soft perspective-aware spatial predicates, and composes them with a training-free Pythonic symbolic executor, separating perception from reasoning while preserving uncertainty through multi-hop inference. We also introduce 3D FORCE, a diagnostic benchmark that controls reasoning depth, view, and perspective composition across spatial arrangement grounding (SAG) and referring expression grounding (REF). On 3D FORCE, VLMs and spatially trained models degrade sharply as depth and perspective complexity increase, whereas SATURN remains stable and outperforms strong baselines. On the real-world MindCube benchmark, SATURN achieves 78.57% overall accuracy, outperforming the strongest baseline by 14 pp.

Danial Kamali, Tanawan Premsri, Shreya Rajpal, Amir Zadeh, Chuan Li, Parisa Kordjamshidi• 2026

Related benchmarks

TaskDatasetResultRank
Spatial Reasoning3D FORCE SAG 1.0 (test)
SAG Score94.4
17
Visual Grounding3D FORCE REF 1.0 (test)
REF Score96.79
17
Multi-view spatial reasoningMMSI
Overall Score48.77
12
Showing 3 of 3 rows

Other info

Follow for update