Share your thoughts, 1 month free Claude Pro on us
See more
Home
/
Benchmarks
Multi-domain Knowledge and Reasoning on MMLU-Pro
Loading...
77.7
Accuracy
Qwen3
19.2936
34.4568
49.62
64.7832
Apr 2, 2026
Apr 4, 2026
Apr 7, 2026
Apr 10, 2026
Apr 13, 2026
Apr 16, 2026
Apr 19, 2026
Accuracy
Average Output Tokens
Updated 3mo ago
Evaluation Results
Method
Method
Links
Accuracy
Average Output Tokens
Qwen3
Size=14B
2026.04
77.7
2,400
Apriel-Reasoner (Ours)
Size=15B
2026.04
77.3
1,900
Phi-4-reasoning
Size=14B
2026.04
77.1
3,400
Nemotron-Cascade
Size=14B
2026.04
76.8
3,600
Apriel-Base
Size=15B
2026.04
76.4
3,500
Apriel-Base + RLVR w/ LP
Size=15B, Length Penal...
2026.04
75.6
1,500
COACT
Trained on=WebInstruct...
2026.04
24.17
-
Pref + Ent
Trained on=WebInstruct...
2026.04
23.45
-
Entropy
Trained on=WebInstruct...
2026.04
22.91
-
Pref Certainty
Trained on=WebInstruct...
2026.04
22.18
-
Random
Trained on=WebInstruct...
2026.04
21.54
-
Feedback
Search any
task
Search any
task