SoK: Reconstruction Attacks on Synthetic Tabular Data (Insights from Winning the NIST CRC)
About
Synthetic data is increasingly promoted as a privacy-preserving substitute for releasing sensitive tabular records, yet its central adversarial threat ("reconstruction", the recovery of an individual's hidden attribute values from a synthetic release and a handful of known quasi-identifiers) has been studied only in scattered, hard-to-compare settings. We present the first systematization of reconstruction (equivalently, attribute inference) attacks on de-identified and synthetic tabular data. We contribute a taxonomy that organizes attacks by the structure they exploit; the most systematic empirical evaluation to date, pitting fourteen attacks against nine synthetic data generation (SDG) methods across five benchmark datasets; and a set of new attacks that fill gaps in the taxonomy, one of which (CoBP-RA) is the strongest attack we measure. Crucially, we introduce a methodology for interpreting what attack success means: a memorization test that distinguishes reconstruction of the population distribution from memorization of training records, and a reduction that places reconstruction and membership inference on a single comparable scale. Our findings: the choice of SDG method governs risk far more than the choice of attack; differential privacy protects mainly at small budgets ($\varepsilon\lesssim1$), above which protection plateaus, bounded by the synthesizer's capacity rather than its noise; de-identification methods are the most exposed; and most reconstruction reflects distributional structure rather than memorization, concentrating individual risk on atypical records. The attacks and infrastructure are externally validated by our first-place finish among all red teams in the 2025 \textit{National Institute of Standards and Technology} (NIST) Collaborative Research Cycle.
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Reconstruction | CDC Diabetes 1k rows QIdemo | R_adv81.9 | 90 | |
| Feature Reconstruction | Adult 10k rows QIdemo RankSwap release 1.0 (test) | Mean Reconstruction Accuracy (R_adv)26.7 | 19 | |
| Membership Inference Attack | NIST Arizona QImedium | AUC50 | 18 | |
| Membership Inference Attack | CDC Diabetes QIdemo | AUC0.77 | 18 | |
| Membership Inference Attack | Adult QIdemo | AUC65 | 18 | |
| Attribute Reconstruction | Adult 1k | Income R^tr_adv67.8 | 7 | |
| Attribute Reconstruction | Adult 10k | Income R^tr_adv Accuracy70.5 | 7 | |
| Attribute Reconstruction | Arizona 10k | Sex Reconstruction Score ($R^{tr}_{adv}$)70.8 | 7 | |
| Attribute Reconstruction | CDC 1k | Diabetes R^tr_adv58.2 | 7 | |
| Reconstruction Attack | NIST SBO MST, ε=10 QI1 (1k rows) | Reconstruction Accuracy (Radv)32.9 | 5 |