Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Multi-scale Object-Aware Gaze Estimation via Geometric Reasoning

About

Gaze target estimation aims to predict the semantic object an observer fixates upon within an image, a task deeply rooted in the object-oriented nature of human gaze. Observers tend to select a specific semantic entity as the attentional target, rather than responding randomly across arbitrary regions of the image. However, existing methods typically model this task as a direct mapping from global features to gaze heatmaps, essentially treating it as a pixel-level regression problem. This approach fails to explicitly represent the gazed object as a distinct entity, making it difficult to produce stable and semantically consistent predictions in complex scenes. To address this, we propose a two-stage gaze estimation framework guided by object semantics, reformulating gaze target estimation as a hierarchical reasoning process. Our method incorporates object-level representations during feature encoding to align image features with discrete semantic entities, then introduces multi-scale feature fusion and geometric constraints from head pose and gaze direction for fine-grained localization and object-level discrimination. Extensive experiments on GazeFollow, VideoAttentionTarget, ChildPlay, and GOO-Real demonstrate that our method achieves AUC of 0.961, 0.948, 0.987, and 0.977 respectively, delivering strong performance across all benchmarks while maintaining a compact parameter size of 7.1M.

Jiajie Mi, Xinyu Liu, Mengke Song, Chenglizhao Chen• 2026

Related benchmarks

TaskDatasetResultRank
Gaze target estimationGazeFollow
Avg L2 Distance0.094
67
Gaze target estimationVideoAttentionTarget
L2 Distance0.095
56
Gaze target estimationChildPlay
L2 Distance0.084
11
Gaze target estimationGOO-Real
AUC0.977
6
Showing 4 of 4 rows

Other info

Follow for update