Share your thoughts, 1 month free Claude Pro on us
See more
Feedback
Search any
task
Search any
task
SOTA General Language Understanding benchmarks and papers with code | Wizwand
Share your thoughts, 1 month free Claude Pro on us
See more
Home
/
Tasks
General Language Understanding
Benchmarks
Dataset Name
SOTA Method
Dataset Name
SOTA Method
Metric
Trend
Results
Last Updated
tinyBenchmark
No Steering
Accuracy (ARC)
77.51
91
1mo ago
GLUE
Full FT
Accuracy
92.5
75
2mo ago
LAMBADA, ARC-C, HellaSwag, BoolQ, MMLU, GSM8K Average
BF16-Baseline
Accuracy
74.35
56
1mo ago
GLUE v1 (test dev)
BOMF (Ours)
MNLI
87.86
40
4mo ago
MMLU
Origin
MMLU Score
73.59
39
3mo ago
Standard Downstream Tasks Suite (SciQ, PIQA, WinoGrande, ARC-E, ARC-C, HellaSwag, LogiQA, BoolQ, LAMBADA, MMLU)
ConceptLM
Average Accuracy
48.3
32
4mo ago
NLU Suite (MMLU, SST-2, AGNews, 20News, MNLI, SNLI)
ReLoRA
Average Accuracy
89.9
31
1mo ago
MMLU
Qwen3-4B
MMLU Accuracy
72.45
29
1mo ago
Average
EMoE
Average Accuracy
72.93
26
2mo ago
General LLM Benchmarks (ARC-C, CSQA, HellaSwag, LAMBADA, MMLU, OpenBookQA, PIQA, Winogrande) (test)
Original
ARC-C Accuracy
59.5
22
4mo ago
CMMLU, CEval
A-CHORD
Chinese Score (CMMLU/CEval)
71.66
20
1mo ago
MMLU, MMLU-Pro, GPQA-Diamond
A-CHORD
English Score
55.46
20
1mo ago
General Ability Suite (MMLU, PIQA, ARC-E, ARC-C, BoolQ, WinoGrande, HellaSwag, TruthfulQA)
LRC
MMLU Accuracy
65
20
1mo ago
12-task evaluation suite (test)
Efficient-DLM 8B
Average Score
71.62
20
4mo ago
MMLU
LP-SFT
5-shot Accuracy
78.82
18
18d ago
C-Eval (val)
Qwen-1.5 14B (Teacher)
Accuracy
78.68
18
4mo ago
Held-out capability suite (test)
Base
AIME-2024 Accuracy
62.9
16
2mo ago
Overall LLM Evaluation Suite PiQA, ARC, HellaSwag, WinoGrande, MMLU v1
LLaMA-3-8B-Lizard
Overall Accuracy
74.6
16
3mo ago
General Downstream Tasks Aggregate
PonderLM-2-Pythia-1.4B
Average Accuracy
59.5
16
26d ago
General Ability Suite (C-QA, T-QA, LAM, MMLU, L-Code)
Base
Average Score
48.1
16
4mo ago
KMMLU
Solar Open
Overall Score
73
16
1mo ago
10-benchmark suite average
FP
Average Accuracy
73.81
15
1mo ago
10 Benchmarks Average (test)
Base
Accuracy (Average)
67.4
15
2mo ago
8 Sub-Tasks (test)
LoRA
Performance on 8 Sub-Tasks
62.3
14
4mo ago
NLP Evaluation Suite (SciQ, PIQA, WG, ARC, HellaSwag, LogiQA, BoolQ, LAMBADA)
MOUE L48
SciQ Accuracy
58.3
14
4mo ago
Showing 25 of 70 rows
25 / page
50 / page
100 / page
1
2
3
Search any
task
Search any
task
Privacy Policy
Terms of Service
FAQs
Swarm Docs