Calibrating Data to Sensitivity in Private Data Analysis

About

We present an approach to differentially private computation in which one does not scale up the magnitude of noise for challenging queries, but rather scales down the contributions of challenging records. While scaling down all records uniformly is equivalent to scaling up the noise magnitude, we show that scaling records non-uniformly can result in substantially higher accuracy by bypassing the worst-case requirements of differential privacy for the noise magnitudes. This paper details the data analysis platform wPINQ, which generalizes the Privacy Integrated Query (PINQ) to weighted datasets. Using a few simple operators (including a non-uniformly scaling Join operator) wPINQ can reproduce (and improve) several recent results on graph analysis and introduce new generalizations (e.g., counting triangles with given degrees). We also show how to integrate probabilistic inference techniques to synthesize datasets respecting more complicated (and less easily interpreted) measurements.

Davide Proserpio, Sharon Goldberg, Frank McSherry• 2012

Related benchmarks

Task	Dataset	Result
Image Classification	MNIST	Accuracy97.42	398
Image Classification	Fashion MNIST	Accuracy83.67	317
Regression	Communities and Crime 1990 US Census / 1990 US LEMAS / 1995 FBI UCR (test (20%))	MSE (Mean)0.0182	78
Regression	California Housing Standard (test)	MSE0.5922	78
Regression	Criteo Sponsored Search Conversion Log (test)	MSE3.12e+3	78
Natural Language Inference	QNLI	Accuracy83.26	78
Data-to-text generation	E2E (test)	BLEU23.39	49
Collision Attack	Input Reconstruction (First split)	BERTScore0.973	27
Collision Attack	Input Reconstruction (Mid)	BERTScore0.995	27
Inversion Attack	Input Reconstruction (First split)	BERTScore63.3	27

Showing 10 of 22 rows

Other info

Follow for update

@wizwand_team Discord