Share your thoughts, 1 month free Claude Pro on us
See more
Home
/
Benchmarks
Knowledge-intensive Question Answering on NQ
Loading...
39.8
Truthfulness
GPT-5
-3.256
7.922
19.1
30.278
Sep 30, 2025
Truthfulness
Hallucination
Accuracy
Updated 1mo ago
Evaluation Results
Method
Method
Links
Truthfulness
Hallucination
Accuracy
GPT-5
2025.09
39.8
28
-
OpenAI o3
2025.09
38.1
30.8
-
TruthRL
Backbone=Llama3.3-70B-...
2025.09
30.7
26
-
TruthRL
Retrieval Setting=With...
2025.09
28.8
24.9
53.7
Prompting
Backbone=Llama3.3-70B-...
2025.09
28.1
30.8
-
TruthRL
Retrieval Setting=With...
2025.09
26.4
21.2
47.6
TruthRL
Retrieval Setting=With...
2025.09
12.9
30.9
43.8
TruthRL
Retrieval Setting=With...
2025.09
-1.6
25
23.5
Feedback
Search any
task
Search any
task