Share your thoughts, 1 month free Claude Pro on us
See more
Feedback
Search any
task
Search any
task
SOTA Model Merging benchmarks and papers with code | Wizwand
Share your thoughts, 1 month free Claude Pro on us
See more
Home
/
Tasks
Model Merging
Benchmarks
Dataset Name
SOTA Method
Dataset Name
SOTA Method
Metric
Trend
Results
Last Updated
8 Vision tasks (test)
Individual Fine-tuned
Accuracy
95.8
77
2mo ago
Average of 8 benchmarks
Pico
Average Accuracy
52.79
72
3mo ago
GLUE CoLA, MRPC, RTE, SST-2
Task arithmetic
Absolute Accuracy
75.9
60
4mo ago
Llama-3.2-1B-Instruct Task Set
K-Merge
Score S(gamma)
0.83
56
23d ago
20-task vision merging scenario (test)
Individual Fine-tuned
Accuracy
94.7
44
2mo ago
14-task vision merging scenario (test)
Individual Fine-tuned
Accuracy
94.3
44
2mo ago
Language Benchmarks 5-task
Indiv.
Score
0.63
44
2mo ago
Large-scale tasks
DOGE AM
Average Normalized Accuracy
98.2
36
2mo ago
TED Talks and XLSum 40 sequential tasks (test)
K-Merge
S^(γ)
84
27
23d ago
Sustainability to large-scale tasks
DOGE AM
Average Normalized Accuracy
91.4
24
2mo ago
Sustainability to large-scale tasks 2 tasks
DOGE AM
Average Normalized Accuracy
101.2
24
2mo ago
7 NLP tasks (test)
EXPERTS
Accuracy
79.2
22
3mo ago
7-task NLP benchmark
METIS
Avg Performance
1.18
20
1mo ago
8 Vision Tasks (average)
PACT-Iso-C
Average Accuracy
86.3
18
1mo ago
Large-scale tasks 16 tasks merged
DOGE AM
Average Normalized Accuracy
91.5
12
2mo ago
Large-scale tasks 12 tasks merged
DOGE AM
Avg Normalized Acc
94.3
12
2mo ago
Sustainability to large-scale tasks (20 tasks)
RegMean++
Average Normalized Accuracy
82.9
12
2mo ago
Sustainability to large-scale 8 tasks
DOGE AM
Avg Normalized Accuracy
94.8
12
2mo ago
Sustainability to large-scale tasks 4 tasks
DOGE AM
Average Accuracy
98.3
12
2mo ago
LLM Evaluation Suite
KARCHER
Normalized Score
0.401
12
4mo ago
Vision, Language, and Multi-modal tasks
Multiple Models
Parameters
8
11
2mo ago
model merging benchmark Many-shot
METIS
Average
1.015
9
1mo ago
Multi-task Evaluation Suite Instruction, Math, Multilingual, Safety
METIS
Average Score
1.015
9
1mo ago
CIFAR100, Cars196, SUN397, EuroSAT, GTSRB, Pets
Individual
Clean Accuracy
83.22
7
1mo ago
Qwen3-4B-Base Transfer 8 benchmarks
Pico
Math Accuracy
32.65
6
3mo ago
Showing 25 of 26 rows
25 / page
50 / page
100 / page
1
2
Search any
task
Search any
task
Privacy Policy
Terms of Service
FAQs
Swarm Docs