Share your thoughts, 1 month free Claude Pro on us
See more
Home
/
Benchmarks
Supervised Fine-Tuning on Supervised Fine-Tuning (SFT) Evaluation Tasks
Loading...
92.4
Textual Entailment
TANDEM
91.88
92.015
92.15
92.285
Jun 3, 2026
Textual Entailment
Answer Verification
Text Matching
Inference Extraction
Word Semantics
Text Categorization
Average Metric Score
Test Loss
Updated 1mo ago
Evaluation Results
Method
Method
Links
Textual Entailment
Answer Verification
Text Matching
Inference Extraction
Word Semantics
Text Categorization
Average Metric Score
Test Loss
TANDEM
Model Scale=3B, Evalua...
2026.06
92.4
76.9
89.6
80.1
88
87.9
85.8
0.174
Uniform
Model Scale=3B, Evalua...
2026.06
91.9
74.8
88.8
79.2
88.8
85.5
84.9
0.189
Feedback
Search any
task
Search any
task