Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

EviRank: Evidence-Based Confidence Estimation for LLM-Based Ranking

About

Large Language Models show promise for recommendation, but they raise reliability concerns due to limited domain coverage and inherent stochasticity. Existing uncertainty quantification methods persist two fundamental challenges: (1) the global confidence score designed for question answering fails to reveal which positions are unreliable in ranking list; (2) fine-grained confidence extracted from model internals exhibits uniformly low values across all positions, making it impossible to filter unreliable predictions. To tackle the challenges, we propose an evidence-based confidence estimation for LLM-based ranking (EviRank). We extract three complementary evidences from a single forward pass and aggregate them via reliable opinion aggregation. Furthermore, we recognize that ranking positions are inherently unequal, and introduce a position-aware calibration. Lastly, the calibrated confidence guides ranking optimization. Experiments on three datasets demonstrate that our method achieves state-of-the-art performance on both recommendation and uncertainty quantification.

Meng Yan, Cai Xv, Xujing Wang, Ziyu Guan, Wei Zhao• 2026

Related benchmarks

TaskDatasetResultRank
RecommendationSteam
Recall@565.83
35
Ranking Uncertainty QuantificationMovieLens 1M
Kendall's Tau@50.1576
12
Ranking Uncertainty QuantificationAmazon Grocery
Tau@50.2661
12
Ranking Uncertainty QuantificationSteam
Tau@50.3133
12
RecommendationMovieLens 1M
R@557.02
9
RecommendationAmazon Grocery
R@556.02
9
Showing 6 of 6 rows

Other info

Follow for update