Share your thoughts, 1 month free Claude Pro on us
See more
Feedback
Search any
task
Search any
task
SOTA Generative Question Answering benchmarks and papers with code | Wizwand
Share your thoughts, 1 month free Claude Pro on us
See more
Home
/
Tasks
Generative Question Answering
Benchmarks
Dataset Name
SOTA Method
Dataset Name
SOTA Method
Metric
Trend
Results
Last Updated
Bolmo Evaluation Suite GenQA 7B
Llama 3.1 70B
GenQA Average
81.6
39
3mo ago
MsMARCO (test)
Match-LSTM
ROUGE Score
40.7
18
4mo ago
MsMARCO (dev)
RAG
ROUGE Score
57.2
11
4mo ago
SimpleQA
RCSP
HALL Score
92
10
1mo ago
MedHallu
Base
HALL Score
67.33
10
1mo ago
WebQuestions
Base
HALL Score
52
10
1mo ago
TruthfulQA (test)
Plan-and-Solve
HALL Score
82.33
10
1mo ago
HaluEval (test)
Plan-and-Solve
HALL Rate
50.33
10
1mo ago
Lu Xun's essay collections
CharacterBot
Content Score
3.758
10
4mo ago
SQuAD Clean (test)
SDBN-p
Exact Match (EM)
72.84
8
1mo ago
Amazon (test)
Prior-Aug
EM
57.99
8
4mo ago
Reddit (test)
MLE
EM
61.19
8
4mo ago
BioASQ (test)
SWEP
EM
43.01
8
4mo ago
NYT (test)
SWEP
EM
76.42
8
4mo ago
Wiki (test)
SWEP
EM
73.34
8
4mo ago
FatwaQA
Gemini-3-Pro
Accuracy
67
7
4mo ago
SQuAD Double-Char (test)
SDBN-p
EM
69.52
5
1mo ago
SQuAD Keyboard-Char (test)
SDBN-p
EM
69.04
5
1mo ago
DriveLM (test)
DriveLM-Agent
BLEU-4
53.09
5
4mo ago
TruthfulQA
KLAS
ROUGE-1
64.5
4
1mo ago
SQuAD Del-Word (test)
SDBN
EM
51.7
3
1mo ago
SQuAD Del-Char (test)
SDBN
Exact Match (EM)
54.1
3
1mo ago
SQuAD
Blended RAG
EM
57.63
3
4mo ago
Showing 23 of 23 rows
25 / page
50 / page
100 / page
1
Search any
task
Search any
task
Privacy Policy
Terms of Service
FAQs
Swarm Docs