Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

BYORn: Bootstrap Your Own Responses to Defend Large Vision-Language Models Against Backdoor Attacks

About

Supervised fine-tuning is the predominant approach for adapting autoregressive vision-language models to downstream tasks. Recent work has shown that this paradigm is highly vulnerable to backdoor attacks, and that existing defenses are ineffective in open-ended generation settings. In response, we propose BYORn, a backdoor-robust fine-tuning framework motivated by the observation that poisoned target responses are often semantically implausible given the corresponding image-text inputs and a pretrained model. BYORn identifies such misaligned responses and dynamically replaces them with alternative responses generated by the model, thereby breaking the correlation between triggers and target outputs. The resulting objective gradient corresponds to the gradient of the empirical estimate of the population risk upper bound over the clean data distribution. Empirically, BYORn consistently improves robustness to backdoor attacks while preserving clean-task performance, establishing a new trade-off frontier between generalization and attack success rate. Finally, we demonstrate that BYORn remains effective against adaptive attacks specifically designed to circumvent the proposed defense.

Ivan Saboli\'c, Marin Or\v{s}i\'c, Josip \v{S}ari\'c, Sven Lon\v{c}ari\'c• 2026

Related benchmarks

TaskDatasetResultRank
Image CaptioningMS-COCO (val)
CIDEr90.9
20
Image CaptioningFlickr30k (val)
CIDEr62
20
Spot the DifferenceCGD (val)
CIDEr149.4
10
Visual Question AnsweringScienceQA
Accuracy (ACC)88.9
10
Showing 4 of 4 rows

Other info

Follow for update