Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Modeling Human Gaze Behavior with Diffusion Models for Unified Scanpath Prediction

About

Predicting human gaze scanpaths is crucial for understanding visual attention, with applications in human-computer interaction, autonomous systems, and cognitive robotics. While deep learning models have advanced scanpath prediction, most existing approaches generate averaged behaviors, failing to capture the variability of human visual exploration. In this work, we present ScanDiff, a novel architecture that combines diffusion models with Vision Transformers to generate diverse and realistic scanpaths. Our method explicitly models scanpath variability by leveraging the stochastic nature of diffusion models, producing a wide range of plausible gaze trajectories. Additionally, we introduce textual conditioning to enable task-driven scanpath generation, allowing the model to adapt to different visual search objectives. Experiments on benchmark datasets show that ScanDiff surpasses state-of-the-art methods in both free-viewing and task-driven scenarios, producing more diverse and accurate scanpaths. These results highlight its ability to better capture the complexity of human visual behavior, pushing forward gaze prediction research. Source code and models are publicly available at https://aimagelab.github.io/ScanDiff.

Giuseppe Cartella, Vittorio Cuculo, Alessandro D'Amelio, Marcella Cornia, Giuseppe Boccignone, Rita Cucchiara• 2025

Related benchmarks

TaskDatasetResultRank
Scanpath GenerationCOCO-Search18 Target Absent
LD0.022
18
Scanpath GenerationCOCO-FV (COCO-FreeView)
LD70.118
14
Scanpath GenerationMIT-FV MIT1003
LD35.347
14
Scanpath GenerationCOCO-Search18 Target Present
LD11.299
12
Scanpath PredictionCOCO-FreeView
LD0.1
7
Scanpath PredictionMIT-FV MIT1003
LD0.041
7
Scanpath PredictionCOCO-Search18 Target Present
LD (Likelihood Distance)0.09
6
Showing 7 of 7 rows

Other info

Follow for update