Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Unconditional Truthfulness: Learning Unconditional Uncertainty of Large Language Models

About

Uncertainty quantification (UQ) has emerged as a promising approach for detecting hallucinations and low-quality output of Large Language Models (LLMs). However, obtaining proper uncertainty scores is complicated by the conditional dependency between the generation steps of an autoregressive LLM because it is hard to model it explicitly. Here, we propose to learn this dependency from attention-based features. In particular, we train a regression model that leverages LLM attention maps, probabilities on the current generation step, and recurrently computed uncertainty scores from previously generated tokens. To incorporate the recurrent features, we also suggest a two-staged training procedure. Our experimental evaluation on ten datasets and three LLMs shows that the proposed method is highly effective for selective generation, achieving substantial improvements over rivaling unsupervised and supervised approaches.

Artem Vazhentsev, Ekaterina Fadeeva, Rui Xing, Gleb Kuzmin, Ivan Lazichny, Alexander Panchenko, Preslav Nakov, Timothy Baldwin, Maxim Panov, Artem Shelmanov• 2024

Related benchmarks

TaskDatasetResultRank
Hallucination DetectionTriviaQA--
625
Hallucination DetectionNQ-Open
AUROC0.7682
141
Hallucination DetectionCoQA
AUROC74.86
134
Uncertainty QuantificationAggregated Experimental Datasets (XSum, SamSum, CNN, WMT19, MedQUAD, TruthfulQA, CoQA, SciQ, TriviaQA, MMLU, GSM8k) (test)
Mean Rank1
88
Hallucination DetectionSciQ
AUROC0.6675
80
Claim Verification9-dataset aggregate retrieval-free setting (test)
ROC-AUC65.7
70
Question AnsweringMedQUAD
PRR58.3
66
Selective GenerationXsum
ROC-AUC85.9
66
Selective Generationcnn
ROC-AUC75.8
66
Selective GenerationTruthfulQA
ROC-AUC0.744
66
Showing 10 of 42 rows

Other info

Follow for update