Share your thoughts, 1 month free Claude Pro on us
See more
Feedback
Search any
task
Search any
task
SOTA Dialogue Response Generation benchmarks and papers with code | Wizwand
Share your thoughts, 1 month free Claude Pro on us
See more
Home
/
Tasks
Dialogue Response Generation
Benchmarks
Dataset Name
SOTA Method
Dataset Name
SOTA Method
Metric
Trend
Results
Last Updated
MSC
EventWeave
B-4 Score
35.8
38
3mo ago
Chronicle
ReBotGEN + EventWeave
B-4
33.5
38
3mo ago
Persona-Chat
JGR-R
BLEU-1
53.3
20
4mo ago
Topical-Chat Global
GPT-4o mini
Und
98.5
16
4mo ago
DuRecDial 2.0
LLaMA-3B w/ FF
F1
45.84
14
2mo ago
DuRecDial
LLaMA-3B w/ FF
F1
47.43
14
2mo ago
KEEM (KMSC memories) 1.0 (test)
Llama10.8B tuned Korean
Perplexity
5.47
14
4mo ago
ConvAI2
MoCoRP LLM
F1
22.89
12
4mo ago
MSC Session 3
G-Long
BLEU-1 Score
24.18
10
1mo ago
bAbI Dialogue Task 5 OOV
GLMP
Per-response Accuracy
92
9
4mo ago
bAbI Dialogue Task 4 OOV
Ptr-Unk
Per-response Accuracy
100
9
4mo ago
bAbI Dialogue Task 3 OOV
GLMP
Accuracy (Per Response)
96.7
9
4mo ago
bAbI Dialogue Task 2 OOV
GLMP
Accuracy (Per-response)
100
9
4mo ago
bAbI Dialogue Task 1 OOV
GLMP
Per-response Accuracy
1
9
4mo ago
bAbI Dialogue Task 5
QRN
Per-response Accuracy
99.6
9
4mo ago
bAbI Dialogue Task 4
Ptr-Unk
Per-response Accuracy
100
9
4mo ago
bAbI Dialogue Task 3
GLMP
Accuracy (Per-response)
96.3
9
4mo ago
bAbI Dialogue Task 2
MN
Per-response accuracy
100
9
4mo ago
bAbI Dialogue Task 1
GMN
Per-response Accuracy
100
9
4mo ago
LoCoMo
DialogLM + EventWeave
BLEU-4
29.1
8
3mo ago
KEEM memories 1.0 (test)
Llama10.8B tuned Korean
Perplexity
4.56
7
4mo ago
Dialogue Dataset (test)
Sampling
Adversarial Success
37.2
7
4mo ago
GROWOVER-DIALOGUE (ALL)
RiLM
BLEU Score (Month 9)
4.7
6
4mo ago
GROWOVER-DIALOGUE (CHANGED)
RiLM
BLEU (Month 9)
7.26
6
4mo ago
100 randomly sampled conversational pairs (test)
SaBART
Appropriateness
66.1
6
4mo ago
Showing 25 of 53 rows
25 / page
50 / page
100 / page
1
2
3
Search any
task
Search any
task
Privacy Policy
Terms of Service
FAQs
Swarm Docs