Share your thoughts, 1 month free Claude Pro on us
See more
Home
/
Benchmarks
Knowledge-intensive Question Answering on Average (CRAG, NQ, HotpotQA, MuSiQue)
Loading...
36.8
Truthfulness
GPT-5
21.512
25.481
29.45
33.419
Sep 30, 2025
Truthfulness
Hallucination
Updated 1mo ago
Evaluation Results
Method
Method
Links
Truthfulness
Hallucination
GPT-5
2025.09
36.8
28.3
OpenAI o3
2025.09
34
32.7
TruthRL
Backbone=Llama3.3-70B-...
2025.09
29.9
21
Prompting
Backbone=Llama3.3-70B-...
2025.09
22.1
28.6
Feedback
Search any
task
Search any
task