Our new X account is live! Follow @wizwand_team for updates
WorkDL logo mark

Similarity-as-Evidence: Calibrating Overconfident VLMs for Interpretable and Label-Efficient Medical Active Learning

About

Active Learning (AL) reduces annotation costs in medical imaging by selecting only the most informative samples for labeling, but suffers from cold-start when labeled data are scarce. Vision-Language Models (VLMs) address the cold-start problem via zero-shot predictions, yet their temperature-scaled softmax outputs treat text-image similarities as deterministic scores while ignoring inherent uncertainty, leading to overconfidence. This overconfidence misleads sample selection, wasting annotation budgets on uninformative cases. To overcome these limitations, the Similarity-as-Evidence (SaE) framework calibrates text-image similarities by introducing a Similarity Evidence Head (SEH), which reinterprets the similarity vector as evidence and parameterizes a Dirichlet distribution over labels. In contrast to a standard softmax that enforces confident predictions even under weak signals, the Dirichlet formulation explicitly quantifies lack of evidence (vacuity) and conflicting evidence (dissonance), thereby mitigating overconfidence caused by rigid softmax normalization. Building on this, SaE employs a dual-factor acquisition strategy: high-vacuity samples (e.g., rare diseases) are prioritized in early rounds to ensure coverage, while high-dissonance samples (e.g., ambiguous diagnoses) are prioritized later to refine boundaries, providing clinically interpretable selection rationales. Experiments on ten public medical imaging datasets with a 20% label budget show that SaE attains state-of-the-art macro-averaged accuracy of 82.57%. On the representative BTMRI dataset, SaE also achieves superior calibration, with a negative log-likelihood (NLL) of 0.425.

Zhuofan Xie, Zishan Lin, Jinliang Lin, Jie Qi, Shaohua Hong, Shuo Li• 2026

Related benchmarks

TaskDatasetResultRank
Medical Image ClassificationBUSI--
88
Image ClassificationDermaMNIST
Accuracy80.21
23
Medical Image ClassificationOCTMNIST
Accuracy79.8
19
Image ClassificationKvasir
Mean Accuracy88.58
7
Image ClassificationRetina
Mean Accuracy75.22
7
Image ClassificationLC25000
Mean Accuracy0.9923
7
Image ClassificationCHMNIST
Mean Accuracy91.03
7
Image ClassificationBTMRI
Mean Accuracy93.46
7
Image ClassificationCOVID-QU-Ex
Mean Accuracy89.49
7
Image ClassificationKneeXray
Mean Accuracy49.5
7
Showing 10 of 10 rows

Other info

Follow for update