Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

RAGTruth

Benchmarks

Task NameDataset NameSOTA ResultTrend
Hallucination DetectionRAGTruth (test)
AUROC0.9096
99
Hallucination DetectionRAGTruth
AUROC0.8535
79
Hallucination detectionRAGTruth CNN/DM (subsample)
AUROC0.69
45
Hallucination detectionRAGTruth MS MARCO (subsample)
AUROC0.77
45
Hallucination DetectionRAGTruth RT-QA 1.0 (test)
F1 Score0.7885
33
Hallucination DetectionRAGTruth RT-Summ 1.0 (test)
F1 Score0.6966
30
Hallucination DetectionRAGTruth RT-D2T 1.0 (test)
F1 Score0.7383
30
Hallucination DetectionRAGTruth Llama2-13B (test)
Acc83.33
21
Hallucination DetectionRAGTruth Llama2-7B (test)
Accuracy75.76
21
Hallucination DetectionRAGTruth LLaMA3-8B
Recall78.6
19
Hallucination DetectionRAGTruth LLaMA2-13B
Recall80.68
19
Hallucination DetectionRAGTruth LLaMA2-7B
Recall0.8328
19
Token-level hallucination detectionRAGTruth
AP (Token-level)59.4
18
Answer-level hallucination detectionRAGTruth
AP75.96
18
SummarizationRAGTruth summarization (test)
ROUGE-152
18
Question AnsweringRAGTruth
F1 Score45.89
17
Response-level hallucination detectionRAGTruth (test)
AUC74.5
15
Hallucination DetectionRAGTruth Span-level, leakage-clean protocol
AUC0.702
15
Hallucination DetectionRAGTruth summarization task
Precision77
14
Response-level Hallucination DetectionRAGTruth QA
AUROC91.89
13
Hallucination MitigationRAGTruth
Faithfulness98.4
12
Span-level Hallucination DetectionRagTruth-Avg (test)
F1 Score76.63
12
Grounded Text GenerationRAGTruth
F1 Score33.14
11
GroundednessRagTruth
Kendall's Tau0.57
11
Faithfulness detectionRAGTruth
Accuracy90.3
10
Showing 25 of 43 rows