Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Calibrating Large Language Models with Sample Consistency

About

Accurately gauging the confidence level of Large Language Models' (LLMs) predictions is pivotal for their reliable application. However, LLMs are often uncalibrated inherently and elude conventional calibration techniques due to their proprietary nature and massive scale. In this work, we explore the potential of deriving confidence from the distribution of multiple randomly sampled model generations, via three measures of consistency. We perform an extensive evaluation across various open and closed-source models on nine reasoning datasets. Results show that consistency-based calibration methods outperform existing post-hoc approaches. Meanwhile, we find that factors such as intermediate explanations, model scaling, and larger sample sizes enhance calibration, while instruction-tuning makes calibration more difficult. Moreover, confidence scores obtained from consistency have the potential to enhance model performance. Finally, we offer practical guidance on choosing suitable consistency metrics for calibration, tailored to the characteristics of various LMs.

Qing Lyu, Kumar Shridhar, Chaitanya Malaviya, Li Zhang, Yanai Elazar, Niket Tandon, Marianna Apidianaki, Mrinmaya Sachan, Chris Callison-Burch• 2024

Related benchmarks

TaskDatasetResultRank
Natural Language InferenceMNLI
ECE17.99
32
Paraphrase IdentificationPPDB
ECE23.26
32
Mathematical ReasoningAQUA
ECE7.37
24
ClassificationYahoo
Expected Calibration Error (ECE)18.45
24
Text ClassificationYahoo Answers
ECE21.04
24
CalibrationHellaSwag
Expected Calibration Error (ECE)8.79
22
CalibrationCSQA
ECE12.18
22
CalibrationMSciNLI
Expected Calibration Error (ECE)28.09
22
Confidence calibrationAverage of four domains Relational Inference Planning
Brier Score0.114
18
ClassificationMSciNLI
Expected Calibration Error (ECE)26.07
16
Showing 10 of 25 rows

Other info

Follow for update