Share your thoughts, 1 month free Claude Pro on us
See more
Feedback
Search any
task
Search any
task
SOTA Hallucination Mitigation benchmarks and papers with code | Wizwand
Share your thoughts, 1 month free Claude Pro on us
See more
Home
/
Tasks
Hallucination Mitigation
Benchmarks
Dataset Name
SOTA Method
Dataset Name
SOTA Method
Metric
Trend
Results
Last Updated
KG-FPQ
Retrieval-Augmented Logical Reasoning
Accuracy
95.2
31
4mo ago
CHAIR
MiddleLayer
CHAIR_S
75
24
4mo ago
HallusionBench visual-dependent setting (test)
ReVisiT
qAcc (Question Accuracy)
33.21
21
2mo ago
SHR
AdaVBoost
HSR
22.1
15
4mo ago
Factual Grounding and Causal Reasoning Evaluation Set
CIP
AC
4.25
14
4mo ago
FinLLM-Eval
BALTO
Faithfulness Score
92.5
12
1mo ago
RAGTruth
BALTO
Faithfulness
98.4
12
1mo ago
ConFiQA
BALTO
Faith
99
12
1mo ago
VRIPT-HAL
VideoChat-R1
F1 Score
52.1
12
4mo ago
MME Existence, Count, Position, Color
VAF
Existence Score
195
12
4mo ago
POPE (val)
VIB-Probe
ACC
88.2
10
4mo ago
Hallusion
UniMRG
Accuracy
64.56
10
4mo ago
SHR (test)
DoLa(high)
SPI
5.36
9
3mo ago
Delulu N=1,950, 7 languages (held-out)
Qwen2.5-Coder-7B
Exact Match (EM)
61.8
8
1mo ago
Lyrics Scenario Aggregated
NIM4-ASR
Hallucination Rate
0.081
8
3mo ago
Code-switch Scenario Aggregated
NIM4-ASR
Hallucination rate
0.261
8
3mo ago
English Scenario Aggregated
Qwen3-Omni-Inst
Hallucination Rate
0.7
8
3mo ago
Dialect Scenario Aggregated
NIM4-ASR
Hallucination Rate
11.7
8
3mo ago
Mandarin Scenario Aggregated
NIM4-ASR
Hallucination Rate
0.2
8
3mo ago
UrbanSound8K
Baseline
HR (%)
95.98
6
1mo ago
RAGTruth
SIFT
Hallucination Rate
32.7
6
3mo ago
PAVE
Qwen-VL-Chat
CHAIRi Score
26.78
4
4mo ago
Showing 22 of 22 rows
25 / page
50 / page
100 / page
1
Search any
task
Search any
task
Privacy Policy
Terms of Service
FAQs
Swarm Docs