Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Data Shapley in One Training Run

About

Data Shapley provides a principled framework for attributing data's contribution within machine learning contexts. However, existing approaches require re-training models on different data subsets, which is computationally intensive, foreclosing their application to large-scale models. Furthermore, they produce the same attribution score for any models produced by running the learning algorithm, meaning they cannot perform targeted attribution towards a specific model obtained from a single run of the algorithm. This paper introduces In-Run Data Shapley, which addresses these limitations by offering scalable data attribution for a target model of interest. In its most efficient implementation, our technique incurs negligible additional runtime compared to standard model training. This dramatic efficiency improvement makes it possible to perform data attribution for the foundation model pretraining stage for the first time. We present several case studies that offer fresh insights into pretraining data's contribution and discuss their implications for copyright in generative AI and pretraining data curation.

Jiachen T. Wang, Prateek Mittal, Dawn Song, Ruoxi Jia• 2024

Related benchmarks

TaskDatasetResultRank
Protected-attribute detection-gap evaluationAdult
L^TPR0.7
14
Corruption DetectionImageNet100 label noise
F1 Score100
10
Noisy label detectionCIFAR-10 1k subset 10% label noise
AUC (%)80.28
8
Vision data selectionCIFAR-10 1k 10% label noise
Accuracy (20% Selected Data)55.6
8
Corrupted-sample detectionElectricity
F1-score41
7
Corrupted-sample detectionCIFAR10
F1 Score51
7
Corrupted-sample detection2Dplanes
F1 Score64
7
Corrupted-sample detectionBBC
F1-score62
7
Corrupted-sample detectionIMDB
F1-score45
7
Corrupted-sample detectionSTL10
F1-score59
7
Showing 10 of 15 rows

Other info

Follow for update