Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

TextDS: Parameter-Efficient Representation Alignment for Scene Text Detection under Distribution Shifts

About

In real-world deployments, scene text detectors inevitably face distribution shifts beyond the training distribution. Prior work often depends on large-scale scene-text pretraining, yet evaluation under cross-domain changes and real-world imaging degradations remains limited. We propose TextDS, an efficient framework for scene text detection under distribution shifts. First, we propose a data-efficient dual-encoder design with visual foundation models, eliminating the reliance on large-scale scene-text pretraining. Second, we introduce Step-wise LoRA adaptation (SWLoRA), which performs progressive low-rank refinement with a dynamic early-exit mechanism for effective feature adaptation. Third, we propose Common Subspace Fusion (CSF) to align and fuse the two branches in a shared subspace while retaining complementary, shift-robust information. Finally, we construct adverse-condition scene text detection datasets to address the gap in evaluating under imaging degradation. Experiments show that TextDS achieves competitive performance in scene text detection, demonstrating robustness across domains and adverse imaging conditions with only 4.9M trainable parameters.

Boyuan Chen, Zichen Dang, Chuang Yang, Lap-Pui Chau, Yi Wang• 2026

Related benchmarks

TaskDatasetResultRank
Text DetectionCTW1500
F-measure90.6
108
Scene Text DetectionTotal-Text
Precision89.7
89
Text DetectionMLT
Recall92
10
Showing 3 of 3 rows

Other info

Follow for update