Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Correcting Prompt Dependence in LLM Benchmarks: A Bayesian Hierarchical Model with Embedding-Space Clustering

About

LLM benchmarking metrics often misstate performance and uncertainty as they rely on two assumptions that frequently do not hold in practice: (i) a sufficient number of evaluations are available for classical inference, and (ii) test prompts are independent. We propose a corrective Bayesian hierarchical model with embedding-space clustering that provides robust performance metrics in limited-data settings while correcting for prompt dependence. We apply the approach to adversarial robustness benchmarks, showing consistent recovery of clustering structure, resulting in more reliable performance metrics, with 4-73% improvements to mean absolute errors and 40-450 unit improvements to expected log posterior densities.

Mary Llewellyn, Isobel Thornton, James Bishop, Annie Gray• 2025

Related benchmarks

TaskDatasetResultRank
Vulnerability scanning and safety evaluationREPEAT
ELPD-87.3
12
Vulnerability scanning and safety evaluationJAVASCRIPT
ELPD-54.8
12
Vulnerability scanning and safety evaluationANSIRAW
ELPD-203.1
12
Vulnerability scanning and safety evaluationEn-Fr
ELPD-230.9
12
Showing 4 of 4 rows

Other info

Follow for update