Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Task Singular Vectors: Reducing Task Interference in Model Merging

About

Task Arithmetic has emerged as a simple yet effective method to merge models without additional training. However, by treating entire networks as flat parameter vectors, it overlooks key structural information and is susceptible to task interference. In this paper, we study task vectors at the layer level, focusing on task layer matrices and their singular value decomposition. In particular, we concentrate on the resulting singular vectors, which we refer to as Task Singular Vectors (TSV). Recognizing that layer task matrices are often low-rank, we propose TSV-Compress (TSV-C), a simple procedure that compresses them to 10% of their original size while retaining 99% of accuracy. We further leverage this low-rank space to define a new measure of task interference based on the interaction of singular vectors from different tasks. Building on these findings, we introduce TSV-Merge (TSV-M), a novel model merging approach that combines compression with interference reduction, significantly outperforming existing methods.

Antonio Andrea Gargiulo, Donato Crisostomi, Maria Sofia Bucarelli, Simone Scardapane, Fabrizio Silvestri, Emanuele Rodol\`a• 2024

Related benchmarks

TaskDatasetResultRank
Visual Question AnsweringVizWiz
Accuracy43.73
1820
Science Question AnsweringScienceQA
Accuracy33.72
791
Natural Language InferenceRTE
Accuracy91.3
590
Image ClassificationTinyImageNet (test)
Accuracy76.18
499
ClassificationCars
Accuracy71.2
492
Image ClassificationDTD
Accuracy90
487
Image ClassificationRESISC45
Accuracy90.6
472
Image ClassificationSVHN (test)
Accuracy96.3
470
Image ClassificationSUN397
Accuracy68.2
450
Image ClassificationMNIST
Accuracy85.3
398
Showing 10 of 203 rows
...

Other info

Follow for update