Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Attend to Anything: Foundation Model for Unified Human Attention Modeling

About

Existing human attention (saliency) modeling methods persist as highly fragmented across modalities, scenes, and task formulations. Consequently, even with increasing model capacity and data scale, current models predominantly remain scene-dependent and task-specific, failing to practically generalize in real-world applications. To address the fundamental limitations, we present the Attend to Anything Model (AAM), a multi-modal foundation model that unifies attention modeling across various image, video, and audio-visual tasks and scenes. AAM reformulates attention as a cognitive entailment relationship organized in a general-to-specific hierarchy, implemented through language prompts with hierarchical embeddings in hyperbolic space. Furthermore, to unify static image and dynamic video attention, we adopt a fluid-dynamics perspective, formulating video-frame attention as a diffusive temporal evolution governed by the Fokker--Planck equation. Extensive experiments on 16 benchmarks demonstrate that AAM consistently outperforms state-of-the-art methods by an average of 6\% across various scenarios, while achieving approximately a 4$\times$ speedup in video inference. Overall, these results demonstrate that AAM provides a principled foundation for future research on attention and saliency-related tasks. The dataset and code will be available at https://github.com/wz-zhao/Attend-to-Anything.

Wenzhuo Zhao, Ronghao Xian, Keren Fu, Qijun Zhao• 2026

Related benchmarks

TaskDatasetResultRank
Video saliency predictionHollywood-2 (test)
SIM0.599
88
Video saliency predictionUCF Sports (test)
SIM0.584
76
Saliency PredictionSalECI E-Commercial
CC0.797
21
Saliency PredictionU-EYE Web page
CC0.743
17
Video Attention ModelingDHF1K
AUC0.919
14
Video Attention ModelingUCF
AUC94.3
14
Video saliency predictionHollywood2
AUC-J0.944
14
Audio-Video Attention ModelingETMD
Correlation Coefficient0.655
13
Audio-Video Attention ModelingSumMe
CC0.55
13
Audio-Video Attention ModelingCoutrot1
CC0.626
13
Showing 10 of 23 rows

Other info

Follow for update