Share your thoughts, 1 month free Claude Pro on us
See more
Home
/
Benchmarks
Knowledge-intensive Question Answering on HotpotQA
Loading...
63.6
Truthfulness
OpenAI o3
7.648
22.174
36.7
51.226
Sep 30, 2025
Truthfulness
Hallucination Rate
Accuracy
Updated 1mo ago
Evaluation Results
Method
Method
Links
Truthfulness
Hallucination Rate
Accuracy
OpenAI o3
2025.09
63.6
18
-
GPT-5
2025.09
62.3
17.3
-
TruthRL
Backbone=Llama3.3-70B-...
2025.09
48.5
17.4
-
Prompting
Backbone=Llama3.3-70B-...
2025.09
42.5
19.6
-
TruthRL
Retrieval Setting=With...
2025.09
33.3
10.7
43.9
TruthRL
Retrieval Setting=With...
2025.09
9.8
12.7
22.5
Feedback
Search any
task
Search any
task