Share your thoughts, 1 month free Claude Pro on us
See more
Home
/
Benchmarks
Dialogue Utterance Generation on User Fidelity Experiment Dataset
Loading...
42.7
ROUGE-L
utt-only-14B
18.676
24.913
31.15
37.387
Jun 28, 2026
ROUGE-L
BLEU-4
Length Ratio
Updated 26d ago
Evaluation Results
Method
Method
Links
ROUGE-L
BLEU-4
Length Ratio
utt-only-14B
Parameter Count=14B, T...
2026.06
42.7
15.9
1.03
CogWM-14B
Parameter Count=14B, T...
2026.06
41.7
15
1.01
DS-V4-Pro
2026.06
28.8
6
1.81
Dual-LLM
Backbone=GPT-5.5
2026.06
25.3
5.4
2.12
GPT-5.5
2026.06
23.7
4.3
1.82
Qwen3-14B
Parameter Count=14B
2026.06
19.6
3.3
2.26
Feedback
Search any
task
Search any
task