Share your thoughts, 1 month free Claude Pro on us
See more
Home
/
Benchmarks
Fidelity Evaluation on BDI/E Dialogue Dataset
Loading...
93
B-M
Qwen3-14B
36.84
51.42
66
80.58
Jun 28, 2026
B-M
D-M
I-M
B-1
Dist-1
Updated 26d ago
Evaluation Results
Method
Method
Links
B-M
D-M
I-M
B-1
Dist-1
Qwen3-14B
number of parameters=14B
2026.06
93
80
51
11.2
95.3
GPT-5.5
2026.06
81
79
27
13.3
96.2
Dual-LLM
2026.06
81
59
28
15.8
96
DSv4-Pro
2026.06
79
66
32
17.2
96.5
CogWM-14B
number of parameters=14B
2026.06
39
26
14
32.2
94.8
Feedback
Search any
task
Search any
task