Share your thoughts, 1 month free Claude Pro on us
See more
Feedback
Search any
task
Search any
task
SOTA General Evaluation benchmarks and papers with code | Wizwand
Share your thoughts, 1 month free Claude Pro on us
See more
Home
/
Tasks
General Evaluation
Benchmarks
Dataset Name
SOTA Method
Dataset Name
SOTA Method
Metric
Trend
Results
Last Updated
Aggregate Benchmarks
Gemini 2.5 Pro+ASP
Average Score
93.9
37
2mo ago
AGIEval
DeepSeek-R1 (BF16)
Accuracy
70.22
29
4mo ago
Average Across all benchmarks
BASTION
Speedup
6.9
28
1mo ago
UltraFeedback Aggregate
LP-SFT
Overall Average Score
59.59
18
18d ago
LiveBench
phi-balancing
Accuracy
46.83
15
2mo ago
All Benchmarks
RESMERGE
Overall Average Score
53.74
12
1mo ago
MM-VET
ECSO
REC
39.5
12
4mo ago
Aggregate Suite PIQA, HellaSwag, WinoGrande, ARC-e, ARC-c
Baseline
Average Score
69
10
4mo ago
Reasoning, Knowledge, and Biomedicine combined datasets (test)
Reasoning
Average Score
60.47
9
4mo ago
RWQA
CoVT-7B
Score
71.8
8
1mo ago
Downstream Suite
DCDM (MoE)
Average Score
39.38
8
2mo ago
ChartQA
DeepLatent-RL-7B
Score
86.4
7
1mo ago
Visual Probe-H
DeepLatent-RL-7B
Score
38.7
7
1mo ago
Average Downstream Benchmark Suite
DoGraph
Average Accuracy
37.9
7
3mo ago
Datacomp small (38 tasks)
Final Self-Filtered 30% Data Mix
Average Score
19.7
6
1mo ago
LiveBench 1125
General Teacher
Score
52.1
6
1mo ago
Instruction Tuning Suite (BIG-bench Hard, MMLU, TyDi QA, MGSM)
Flan-PaLM 2 (L)
Average Score
74.1
4
4mo ago
ExpSuite-Static Overall
ExpGraph
Average Score
78.75
2
1mo ago
Showing 18 of 18 rows
25 / page
50 / page
100 / page
1
Search any
task
Search any
task
Privacy Policy
Terms of Service
FAQs
Swarm Docs