Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Model Reconstruction from Model Explanations

About

We show through theory and experiment that gradient-based explanations of a model quickly reveal the model itself. Our results speak to a tension between the desire to keep a proprietary model secret and the ability to offer model explanations. On the theoretical side, we give an algorithm that provably learns a two-layer ReLU network in a setting where the algorithm may query the gradient of the model with respect to chosen inputs. The number of queries is independent of the dimension and nearly optimal in its dependence on the model size. Of interest not only from a learning-theoretic perspective, this result highlights the power of gradients rather than labels as a learning primitive. Complementing our theory, we give effective heuristics for reconstructing models from gradient explanations that are orders of magnitude more query-efficient than reconstruction attacks relying on prediction interfaces.

Smitha Milli, Ludwig Schmidt, Anca D. Dragan, Moritz Hardt• 2018

Related benchmarks

TaskDatasetResultRank
Model StealingAIDS
AUC89.79
27
Model StealingNCI109
AUC73.21
27
Model StealingNCI1
AUC74.25
27
Model StealingMutagenicity
AUC80.83
18
Graph Model StealingHIV
AUC61.73
9
Graph Model StealingTox21
AUC75.16
9
Graph Model StealingBACE
AUC70.34
9
Node ClassificationOGB-arxiv
AUC87.33
9
Graph Classification Model StealingMutagenicity
AUC80.04
9
Node ClassificationPubmed
AUC85.54
9
Showing 10 of 10 rows

Other info

Follow for update