Share your thoughts, 1 month free Claude Pro on us
See more
Home
/
Benchmarks
Hallucination Mitigation on Delulu N=1,950, 7 languages (held-out)
Loading...
61.8
Exact Match (EM)
Qwen2.5-Coder-7B
42.04
47.17
52.3
57.43
Jun 2, 2026
Exact Match (EM)
Error Score (ES)
Updated 1mo ago
Evaluation Results
Method
Method
Links
Exact Match (EM)
Error Score (ES)
Qwen2.5-Coder-7B
Model size=7B, Trainin...
2026.06
61.8
85
Qwen2.5-Coder-7B
Model size=7B, Trainin...
2026.06
61.4
87
Qwen2.5-Coder-3B
Model size=3B, Trainin...
2026.06
59.2
83.5
Qwen2.5-Coder-3B
Model size=3B, Trainin...
2026.06
58.5
85
Qwen2.5-Coder-3B
Model size=3B, Trainin...
2026.06
52
79.8
Qwen2.5-Coder-7B
Model size=7B, Trainin...
2026.06
48.8
77
Qwen2.5-Coder-3B
Model size=3B, Trainin...
2026.06
45.7
72
Qwen2.5-Coder-7B
Model size=7B, Trainin...
2026.06
42.8
65
Feedback
Search any
task
Search any
task