Share your thoughts, 1 month free Claude Pro on us
See more
Feedback
Search any
task
Search any
task
SOTA Prompt Optimization benchmarks and papers with code | Wizwand
Share your thoughts, 1 month free Claude Pro on us
See more
Home
/
Tasks
Prompt Optimization
Benchmarks
Dataset Name
SOTA Method
Dataset Name
SOTA Method
Metric
Trend
Results
Last Updated
XSum
EGE
Hypervolume (HV)
0.1626
72
2mo ago
Prompt Optimization Benchmark
LinGO
Accuracy
69
24
4mo ago
Logical Reasoning, Mathematical Calculation, and Knowledge Intensive tasks Average
MemAPO
Average Performance (%)
70.7
20
4mo ago
VisEval
PromptAgent
Accuracy (Easy)
0.77
10
4mo ago
DABench
PromptBreeder
Acc (Easy)
80
10
4mo ago
ST
SCULPT
Best Score
71.7
8
1mo ago
FF
LLMLingua
Best Score
98.9
8
1mo ago
CJ
SCULPT
Best Score
76.9
8
1mo ago
DQA
SCULPT
Best Score
82.3
8
1mo ago
GoE
WPRO0.5
Best Performance
43
8
1mo ago
BT
CRAFT
Best Score
62
8
1mo ago
HotpotQA, IFBench, HoVer, PUPA, AIME, and LiveBench-Math 2018-2025 (test)
GEPA
HotpotQA Score
69
8
4mo ago
DSG-1K
CRAFT
DSGScore
0.91
7
4mo ago
P2-hard
Maestro
DSGScore
92
7
4mo ago
10-task prompt optimization suite GSM8K MMLU BBH
ReElicit
Average Win/Tie Rate
81
5
2mo ago
product-gen (test)
Bayesian
Accuracy
92.2
5
2mo ago
trip-advisory (test)
COPRO-R
Accuracy
81.1
5
2mo ago
code-explain (test)
COPRO-R
Accuracy
84.2
5
2mo ago
42 LLM benchmarks Aggregate (overall)
System+Task Optimized
Average Score
67.14
5
3mo ago
CB
TRAS
Accuracy
85.7
4
1mo ago
Biosses
TRAS
Accuracy
70.4
4
1mo ago
Penguins
TRAS
Accuracy
68.6
4
1mo ago
Geometric Shapes
TRAS
Accuracy
63.3
4
1mo ago
Causal Judgment
TRAS
Accuracy
64.4
4
1mo ago
GEPA Evaluation Suite Aggregate
LEVI
Aggregate Score
62.02
4
2mo ago
Showing 25 of 32 rows
25 / page
50 / page
100 / page
1
2
Search any
task
Search any
task
Privacy Policy
Terms of Service
FAQs
Swarm Docs