Share your thoughts, 1 month free Claude Pro on us
See more
Home
/
Benchmarks
Chinese Language Generation on Chinese Dataset
Loading...
4.6
Usefulness Score
Bucket-SFT
2.988
3.4065
3.825
4.2435
Jun 11, 2026
Usefulness Score
Answerability Score
Naturalness Score
Average Score
Updated 1mo ago
Evaluation Results
Method
Method
Links
Usefulness Score
Answerability Score
Naturalness Score
Average Score
Bucket-SFT
n=20
2026.06
4.6
4.95
5
4.85
Full-SFT
n=20
2026.06
4.45
4.65
4.55
4.55
BaseLM
n=19
2026.06
3.47
3.89
3.84
3.74
DPO
n=20
2026.06
3.05
3.6
2.95
3.2
Feedback
Search any
task
Search any
task