Our new X account is live! Follow @wizwand_team for updates
WorkDL logo mark

On the Accuracy of Influence Functions for Measuring Group Effects

About

Influence functions estimate the effect of removing a training point on a model without the need to retrain. They are based on a first-order Taylor approximation that is guaranteed to be accurate for sufficiently small changes to the model, and so are commonly used to study the effect of individual points in large datasets. However, we often want to study the effects of large groups of training points, e.g., to diagnose batch effects or apportion credit between different data sources. Removing such large groups can result in significant changes to the model. Are influence functions still accurate in this setting? In this paper, we find that across many different types of groups and for a range of real-world datasets, the predicted effect (using influence functions) of a group correlates surprisingly well with its actual effect, even if the absolute and relative errors are large. Our theoretical analysis shows that such strong correlation arises only under certain settings and need not hold in general, indicating that real-world datasets have particular properties that allow the influence approximation to be accurate.

Pang Wei Koh, Kai-Siang Ang, Hubert H. K. Teo, Percy Liang• 2019

Related benchmarks

TaskDatasetResultRank
Text ClassificationYelp (test)--
55
End model evaluationCensus
Test Loss0.359
22
End model evaluationDN infograph
Test Loss1.191
22
End model evaluationPW
Test Loss0.299
22
End model evaluationYouTube
Test Loss0.278
22
End model evaluationDN clipart
Test Loss0.872
22
End model evaluationYelp
Test Loss0.425
22
End model evaluationMushroom
Test Loss0.202
22
End model evaluationIMDB
Test Loss0.582
22
End model evaluationDN real
Test Loss0.581
22
Showing 10 of 26 rows

Other info

Follow for update