Private Memorization Editing: Turning Memorization into a Defense to Strengthen Data Privacy in Large Language Models

About

Large Language Models (LLMs) memorize, and thus, among huge amounts of uncontrolled data, may memorize Personally Identifiable Information (PII), which should not be stored and, consequently, not leaked. In this paper, we introduce Private Memorization Editing (PME), an approach for preventing private data leakage that turns an apparent limitation, that is, the LLMs' memorization ability, into a powerful privacy defense strategy. While attacks against LLMs have been performed exploiting previous knowledge regarding their training data, our approach aims to exploit the same kind of knowledge in order to make a model more robust. We detect a memorized PII and then mitigate the memorization of PII by editing a model knowledge of its training data. We verify that our procedure does not affect the underlying language model while making it more robust against privacy Training Data Extraction attacks. We demonstrate that PME can effectively reduce the number of leaked PII in a number of configurations, in some cases even reducing the accuracy of the privacy attacks to zero.

Elena Sofia Ruzzetti, Giancarlo A. Xompero, Davide Venditti, Fabio Massimo Zanzotto• 2025

Related benchmarks

Task	Dataset	Result
Privacy Editing	TDE Email	Leakage0.00e+0	56
Privacy Editing	TDE URL	Leakage0.00e+0	50
Training Data Extraction	email PII	Leakage0.00e+0	45
Training Data Extraction	phone PII	Leak Count0.00e+0	45
Training Data Extraction	URL PII	Leakage2	45
Reliability of post-edit LLMs	Books3	BLEU0.957	36
Reliability of post-edit LLMs	Wikipedia	BLEU97.5	36
Reliability of post-edit LLMs	Pile-CC	BLEU95.1	36
Targeted Data Extraction (TDE) Attack	email PII	PII Leaked Count0.00e+0	21
Targeted Data Extraction (TDE) Attack	phone PII	Leaked Items Count0.00e+0	21

Showing 10 of 18 rows

Other info

Code

Follow for update

@wizwand_team Discord