Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

PAGE: Towards Practical Human-level Gaze Target Estimation

About

Gaze target estimation, the task of predicting where a person is looking in a scene, is crucial to understanding human attention and intent. It is a challenging task that combines high-level understanding of global scene semantics and precise spatial reasoning using human appearance (e.g. pose, eye orientation). As a result, human-level performance remains elusive for existing models, limiting their practical application. To this end, we propose PaGE (Practical Gaze Estimator), a gaze estimation model that explicitly models the complex interaction between scene and head features. Using a PaGE model with a large ViT-H+ backbone as the teacher, we further distill student models with lighter backbones on a much larger and more diverse unlabeled dataset. The architectural improvements and novel training recipe allow PaGE to achieve state-of-the-art performance on several gaze estimation tasks, outperforming humans in 7 out of 9 metrics while reducing the human-AI gap by at least 60% in the remaining 2. The distilled student models retain most of the teacher's performance while being lightweight enough for practical deployment on robots and consumer devices. The code and model checkpoints are available at our project page.

Zhoutong Ye, Chengwen Zhang, Zhaibin Cui, Mingze Sun, Jiaqi Liu, Xiangwu Li, Qingyang Wan, Chang Liu, Xutong Wang, Huan-ang Gao, Yu Mei, Chun Yu, Yuanchun Shi• 2026

Related benchmarks

TaskDatasetResultRank
Gaze target estimationGazeFollow
Avg L2 Distance0.08
67
Gaze target estimationVideoAttentionTarget
L2 Distance0.064
56
Gaze FollowingGOO-Real
AUC93
7
Gaze FollowingZhang
Accuracy76.29
5
Showing 4 of 4 rows

Other info

Follow for update