| Task Name | Dataset Name | SOTA Result | Trend | |
|---|---|---|---|---|
| Zero-shot performance evaluation | LM Eval Harness (HellaSwag, BoolQ, WinoGrande, PiQA, ARC-easy, ARC-challenge) zero-shot | Mean Accuracy75.46 | 60 | |
| Zero-shot Question Answering and Reasoning | LM-Eval-Harness Suite (PIQA, HellaSwag, LAMBADA, ARC-e, ARC-c, SciQ, Race, MMLU) zero-shot | PIQA80.7 | 32 | |
| Common-sense Reasoning | lm-eval-harness 11 tasks | 0-shot Accuracy63.43 | 24 | |
| Commonsense Reasoning | LM-Eval-Harness OBQA, PIQA, ARC-e, ARC-c, HellaSwag, WinoGrande | OBQA Accuracy44.2 | 14 | |
| Zero-shot Evaluation | LM-Eval-Harness | LAMBADA Perplexity (PPL)17.8 | 10 | |
| Multiple Choice Question Answering | LM-Eval-Harness MMLU, ARC-Easy, HellaSwag, PIQA, OpenBookQA, WinoGrande | MMLU Accuracy24.3 | 10 | |
| Question Answering and Commonsense Reasoning | lm-eval-harness PIQA, COPA, OpenBookQA, Winogrande, SciQA, ARC-E, ARC-C | PIQA Accuracy78.8 | 10 | |
| Language Model Evaluation | lm-eval-harness (test) | MMLU74.22 | 9 | |
| Multiple-choice Question Answering and Reasoning | EleutherAI lm-eval-harness ARC-C, ARC-E, HellaSwag, PIQA, WinoGrande standard (test val) | ARC-Challenge Accuracy22.7 | 6 | |
| General Language Model Reasoning | LM-Eval-Harness Hungarian | Arc (hu) Acc38.6 | 4 | |
| Language Modeling Utility | LM Eval Harness | HellaSwag Accuracy0.48 | 3 |