Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Task-Instructed Causal Routing of Vision Foundation Models for Multi-Task Learning

About

Vision foundation models (VFMs) have demonstrated strong robustness and transferability across a wide range of visual tasks. However, each model typically encodes strong inductive biases shaped by its pre-training objective and data domain, resulting in fragmented yet complementary visual knowledge. As a result, a single model often struggles to capture the diverse visual representations required across multiple dense prediction tasks. To address this limitation, we propose TIGER (Task-Instruction-Guided Expert Routing), a framework that coordinates multiple heterogeneous VFMs for multi-task dense prediction. Instead of naively aggregating expert features, TIGER leverages natural-language task instructions to guide a routing network that assigns token-level expert weights conditioned on task semantics, enabling adaptive integration of complementary expert features. TIGER further introduces a counterfactual loss that aligns routing decisions with each expert's causal contribution by measuring prediction changes when experts are excluded, encouraging more reliable and interpretable routing. We evaluate TIGER on two multi-task dense prediction benchmarks, NYUD-v2 and Pascal Context, where it consistently outperforms recent multi-task learning baselines while keeping all VFMs frozen. These results demonstrate that combining instruction-guided expert routing with counterfactual causal alignment enables effective coordination of heterogeneous vision foundation models.

Donghyun Han, Yuseok Bae, Jung Uk Kim, Hyung-Il Kim• 2026

Related benchmarks

TaskDatasetResultRank
Depth EstimationNYU V2
RMSE0.4115
207
Semantic segmentationNYUD v2
mIoU63.55
169
Surface Normal EstimationPascal Context
Mean Error (MAE)12.46
64
Saliency DetectionPascal Context
maxF Score83.97
64
Semantic segmentationPascal Context
mIoU84.58
61
Human ParsingPascal Context
mIoU77.56
54
Boundary DetectionNYUD v2
ODS F-measure80.31
48
Surface Normals EstimationNYUD v2
Mean Error16.8
19
Showing 8 of 8 rows

Other info

Follow for update