Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

LLM Evaluation on Suite of Benchmarks (MMLU, IFEval, GSM8K, MATH, HumanEval, MBPP, Hellaswag, GPQA)

51.94MMLU

Alpaca-GPT4

22.476830.125937.77545.4241Aug 29, 2025
Updated 23d ago

Evaluation Results

MethodLinks
2025.08
51.9438.6850.8710.2817.0743.663.020.5134.5
2025.08
51.3243.251.1812.9239.6341.858.7816.6739.63
2025.08
47.4641.0935.634.9639.6337.448.115.5632.48
2025.08
47.2143.9243.94.229.2743.460.175.5634.7
2025.08
46.1246.1453.312.7240.244853.0512.1238.96
2025.08
44.6947.9657.6218.552.4445.457.3719.742.96
2025.08
43.4740.7865.215.5851.8347.658.6517.6842.6
2025.08
41.545.6660.820.0646.344855.0124.7542.77
2025.08
39.9637.844.55.3840.854442.3827.2735.27
2025.08
37.7134.3553.68119.1545.657.812.5331.48
2025.08
33.9848.8243.826.0635.9842.444.518.1834.22
2025.08
25.5114.7556.3316.5613.4145.625.83024.75
2025.08
23.6129.6143.148.2832.3232.641.83026.42