Share your thoughts, 1 month free Claude Pro on us
See more
Home
/
Benchmarks
LLM-as-a-judge Evaluation on Chinese (test)
Loading...
82.7
Overall Score
Bucket-SFT
47.444
56.597
65.75
74.903
Jun 11, 2026
Overall Score
Utility Score
Conditional Naturalness Score
Distribution Faithfulness Score
Updated 1mo ago
Evaluation Results
Method
Method
Links
Overall Score
Utility Score
Conditional Naturalness Score
Distribution Faithfulness Score
Bucket-SFT
Backbone=Qwen2.5-1.5B,...
2026.06
82.7
80.2
84.8
82.8
Full-SFT
Backbone=Qwen2.5-1.5B,...
2026.06
82.5
79.9
84.7
82.6
Bucket-SFT
Backbone=Llama-3.2-3B,...
2026.06
79.5
76.9
81.8
79.6
BaseLM
Backbone=Llama-3.2-3B,...
2026.06
68.3
66.1
70.1
67.8
BaseLM
Backbone=Qwen2.5-1.5B,...
2026.06
65.2
62.8
67
64.7
HDPO
Backbone=Qwen2.5-1.5B,...
2026.06
49.3
45.9
50
45.9
HDPO
Backbone=Llama-3.2-3B,...
2026.06
48.8
45.8
49.1
45.2
Feedback
Search any
task
Search any
task