Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Can We Infer Confidential Properties of Training Data from LLMs?

About

Large language models (LLMs) are increasingly fine-tuned on domain-specific datasets to support applications in fields such as healthcare, finance, and law. These fine-tuning datasets often have sensitive and confidential dataset-level properties -- such as patient demographics or disease prevalence -- that are not intended to be revealed. While prior work has studied property inference attacks on discriminative models (e.g., image classification models) and generative models (e.g., GANs for image data), it remains unclear if such attacks transfer to LLMs. In this work, we introduce PropInfer, a benchmark task for evaluating property inference in LLMs under two fine-tuning paradigms: question-answering and chat-completion. Built on the ChatDoctor dataset, our benchmark includes a range of property types and task configurations. We further propose two tailored attacks: a prompt-based generation attack and a shadow-model attack leveraging word frequency signals. Empirical evaluations across multiple pretrained LLMs show the success of our attacks, revealing a previously unrecognized vulnerability in LLMs.

Pengrun Huang, Chhavi Yadav, Kamalika Chaudhuri, Ruihan Wu• 2025

Related benchmarks

TaskDatasetResultRank
Property Inference DefenseChatDoctor CC
Generation Attack MAE0.0919
6
Property Inference DefenseChatDoctor QA
Generation Attack MAE0.1301
6
Property Inference DefenseMedCalc QA
Generation Attack MAE0.0653
6
Property Inference DefenseMedCalc CC
Generation Attack MAE0.0264
6
QAChatDoctor
MAE0.0396
5
QAMedCalc
MAE0.0653
5
CCChatDoctor
MAE0.0719
5
CCMedCalc
MAE0.0253
5
Showing 8 of 8 rows

Other info

Follow for update