Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Pre-Generation Hallucination Detection in Large Language Models via Soft-Target Attention Probing

About

Detecting hallucination risk before generation enables abstention, retrieval augmentation, and routing decisions without incurring the cost of decoding. While prior work has shown that such risk can be estimated from a model's internal representations, existing approaches treat this as binary classification over a single decoded output. We instead formulate it as a risk-estimation problem. Under this formulation, we introduce soft-target supervision based on the empirical answer error rate over stochastically sampled outputs - an estimator we prove to be the unique unbiased minimum-variance estimator of the model's per-prompt error probability under its sampling distribution. We further adapt attention probing to the pre-generation setting, enabling the detector to selectively aggregate hallucination-relevant prompt representations. Across three question-answering benchmarks and five models, attention probing outperforms linear probing on short-answer tasks. Replacing binary labels with soft-target supervision further and consistently improves detection quality.

Amina Miftakhova, Alexey Zaytsev• 2026

Related benchmarks

TaskDatasetResultRank
Hallucination DetectionHotpotQA
AUROC0.7484
294
Hallucination DetectionNQ
AUC0.8152
199
Hallucination DetectionSQuAD
AUROC0.8754
127
Hallucination DetectionRAGTruth
AUROC0.6734
79
Hallucination DetectionMMLU-Pro
AUROC75.42
31
Hallucination DetectionBoolQ
ROC-AUC94.46
28
Showing 6 of 6 rows

Other info

Follow for update