Share your thoughts, 1 month free Claude Pro on us
See more
Home
/
Benchmarks
Supervised Fine-Tuning on SFT Mixture 99 tasks (test)
Loading...
85.3
Textual Entailment
DoReMi
81.244
82.297
83.35
84.403
Jun 3, 2026
Textual Entailment
Answer Verification
Text Matching
Information Extraction
Word Semantics
Text Categorization
Test Loss
Average Metric
Updated 1mo ago
Evaluation Results
Method
Method
Links
Textual Entailment
Answer Verification
Text Matching
Information Extraction
Word Semantics
Text Categorization
Test Loss
Average Metric
DoReMi
Model=500M Qwen-2
2026.06
85.3
74.9
83.6
78.1
87
79
0.249
81.6
TANDEM
Model=500M Qwen-2
2026.06
85.3
76.3
86.2
78.6
88.5
84.9
0.208
83.3
Uniform
Model=500M Qwen-2
2026.06
84.4
75.1
86.2
77.8
88.3
83.2
0.231
82.5
Skill-It
Model=500M Qwen-2
2026.06
84.1
75
86.4
78.3
87.9
83.7
0.232
82.6
Aioli
Model=500M Qwen-2
2026.06
83.9
75.8
86.2
78
88.3
84
0.229
82.7
DoGE
Model=500M Qwen-2
2026.06
81.4
76.7
86.5
78.6
88.3
82.4
0.297
82.4
Feedback
Search any
task
Search any
task