Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

SUMO: Segment and Track Any Motion with Nonlinear State Space Models

About

Visual Object Tracking (VOT) and Moving Object Segmentation (MOS) are two fundamental tasks in computer vision that involve both spatial and temporal object dynamics. Existing methods rely predominantly on visual cues and thus often falter in real-world scenarios where object motions are inherently complex and nonlinear. To address this limitation, we propose SUMO, a zero-shot, training-free, unified framework integrating nonlinear dynamics with vision-based segmentation for accurate and consistent VOT and MOS. Specifically, we develop a nonlinear State Space Model (SSM) inspired by robotics principles to capture the complex object dynamics. Building on this model, we propose a Selective Unscented Filter (SUF) for accurate state estimation, which features a joint scoring mechanism and dynamically fuses multi-source predictions to identify the most plausible object state over time. Furthermore, we apply a memory selection mechanism to evaluate the reliability of memory frames. Our extensive experimental results show that SUMO achieves state-of-the-art performance on both VOT and MOS tasks.

Kexin Tian, Sixu Li, Keshu Wu, Yang Zhou, Zhengzhong Tu• 2026

Related benchmarks

TaskDatasetResultRank
Object TrackingLaSoT
AUC74.4
519
Visual Object TrackingGOT-10k
AO83.5
357
Moving Object SegmentationDAVIS Moving 2016
Jaccard Index90.7
48
Video Object SegmentationSegTrack v2
Jaccard Index88.4
42
Moving Object SegmentationFBMS-59
J-Measure90.1
31
Visual TrackingLaSOT ext
AUC61.2
29
Showing 6 of 6 rows

Other info

Follow for update