Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

RaysUp: Ultra-light Universal Feature Upsampling via Geometry-Aware Ray Representation

About

Pre-trained Vision Foundation Models (VFMs) have become central to modern computer vision due to their powerful semantic representations and strong generalization ability. However, their patchified or pooled outputs are inherently low-resolution, limiting their effectiveness in tasks requiring fine-grained, pixel-level reasoning. Existing feature upsampling approaches either degrade semantic fidelity or rely on VFM-specific retraining and heavy architectures, hindering efficiency and scalability. To address these challenges, we propose RaysUp, an ultra-lightweight, task-agnostic, and VFM-agnostic feature upsampling framework that reconstructs high-resolution feature maps at arbitrary resolutions. Unlike conventional 2D interpolation or attention-based schemes, RaysUp lifts feature reconstruction into a geometry-aware ray domain. Specifically, we introduce a Spatially Decoupled Guidance Encoder for direction-aware guidance encoding, an Any-Resolution Cross-Attention mechanism for resolution-flexible reconstruction, and a novel Ray Positional Encoding (RayPE) that injects implicit 3D geometric priors via 6D Plucker ray coordinates. Finally, a Geometry-Aware Neighborhood Attention module further ensures content-adaptive bilateral aggregation while preserving geometric consistency. Extensive experiments across diverse dense prediction tasks demonstrate that RaysUp achieves state-of-the-art performance while using only 16% of the parameters of AnyUp and delivering approximately 7x faster inference. These results highlight a substantially improved accuracy-efficiency trade-off and establish RaysUp as a practical and scalable solution for universal feature upsampling. Code is available at https://github.com/MAP-RaysUp/RaysUp.

Yuchuan Ding, Linfei Li, Lin Zhang, Ying Shen• 2026

Related benchmarks

TaskDatasetResultRank
Semantic segmentationCityscapes
mIoU61.88
674
Semantic segmentationCOCO Stuff
mIoU62.32
421
Semantic segmentationPascal VOC
mIoU84.81
214
Depth EstimationNYU V2
RMSE0.4658
207
Video Object SegmentationDAVIS
J & F Mean71.47
134
Surface Normal EstimationNYU V2--
96
Open Vocabulary Semantic SegmentationCOCO Stuff
mIoU27.11
84
Open-Vocabulary SegmentationCityscapes
mIoU37.24
55
Open-Vocabulary SegmentationADE20K
mIoU20.54
28
Open-Vocabulary SegmentationPascal VOC
mIoU63.28
22
Showing 10 of 11 rows

Other info

GitHub

Follow for update