Share your thoughts, 1 month free Claude Pro on us
See more
Home
/
Benchmarks
Multi-task Language Understanding on MMLU (Subject Accuracy Breakdown)
Loading...
45.5
Average Accuracy
LLAMA-1-7B
24.492
29.946
35.4
40.854
Apr 21, 2026
May 4, 2026
May 17, 2026
May 30, 2026
Jun 12, 2026
Jun 25, 2026
Jul 8, 2026
Average Accuracy
STEM Accuracy
Humanities Accuracy
Social Science Accuracy
Others Accuracy
Updated 16d ago
Evaluation Results
Method
Method
Links
Average Accuracy
STEM Accuracy
Humanities Accuracy
Social Science Accuracy
Others Accuracy
LLAMA-1-7B
Bits=FP16
2026.04
45.5
36.1
43.3
51.6
51.8
Dense
Model=LLaMA-2-7B, Spar...
2026.07
45.2
-
-
-
-
PALS
Model=LLaMA-2-7B, Spar...
2026.07
42.2
-
-
-
-
Wanda
Model=LLaMA-2-7B, Spar...
2026.07
41.5
-
-
-
-
Magnitude
Model=LLaMA-2-7B, Spar...
2026.07
36.4
-
-
-
-
LBLLM
Bits=W(1+1)A4, Backbon...
2026.04
28.8
26.4
28.8
27.7
32
ARB-LLM
Bits=W(1+1)A4, Backbon...
2026.04
27.7
30.2
25.4
30.5
26
Atom
Bits=W(1+1)A4, Backbon...
2026.04
25.3
25.9
24.9
24
26.3
Feedback
Search any
task
Search any
task