Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Thermometer: Towards Universal Calibration for Large Language Models

About

We consider the issue of calibration in large language models (LLM). Recent studies have found that common interventions such as instruction tuning often result in poorly calibrated LLMs. Although calibration is well-explored in traditional applications, calibrating LLMs is uniquely challenging. These challenges stem as much from the severe computational requirements of LLMs as from their versatility, which allows them to be applied to diverse tasks. Addressing these challenges, we propose THERMOMETER, a calibration approach tailored to LLMs. THERMOMETER learns an auxiliary model, given data from multiple tasks, for calibrating a LLM. It is computationally efficient, preserves the accuracy of the LLM, and produces better-calibrated responses for new tasks. Extensive empirical evaluations across various benchmarks demonstrate the effectiveness of the proposed method.

Maohao Shen, Subhro Das, Kristjan Greenewald, Prasanna Sattigeri, Gregory Wornell, Soumya Ghosh• 2024

Related benchmarks

TaskDatasetResultRank
Honesty AlignmentHonestyBench In-Domain
NQ Score58.15
13
Honesty AlignmentHonestyBench OOD
Squad Score60.23
13
Question Answering CalibrationOOD Evaluation (Squad, WQ, CWQ, MSQ, PopQA)
Squad Calibration Score5
11
Question Answering CalibrationIn-Domain Evaluation NQ, TQ, HQ, 2Wiki, Pararel
Calibration Error (NQ)0.06
11
Showing 4 of 4 rows

Other info

Follow for update