Share your thoughts, 1 month free Claude Pro on us
See more
Feedback
Search any
task
Search any
task
SOTA Robustness Evaluation benchmarks and papers with code | Wizwand
Share your thoughts, 1 month free Claude Pro on us
See more
Home
/
Tasks
Robustness Evaluation
Benchmarks
Dataset Name
SOTA Method
Dataset Name
SOTA Method
Metric
Trend
Results
Last Updated
Robust LQR
Pathwise Robust ZO Both
Mean Normalized Performance
0.57
70
1mo ago
Robust LQR adversarial Normalized performance (test)
Pathwise Robust ZO Both
Normalized Performance
1.603
70
1mo ago
SI-Score size synthetic
RobustViT
R@1
66.5
31
4mo ago
SI-Score rotation synthetic
RobustViT
R@1
58
31
4mo ago
SI-Score location synthetic
RobustViT
R@1
48.3
31
4mo ago
MultiRLVR
Master-RM
FPR (%)
0.02
20
4mo ago
MATH
AdvJudge-Zero
FPR (%)
0
20
4mo ago
GSM8K
Master-RM
FPR (%)
0
20
4mo ago
AIME
Master-RM
FPR
0
20
4mo ago
HellaSwag SAGE-generated
GPT-4o
Overall Accuracy (OA)
74.17
12
2mo ago
R-Bench 2024 (test)
Robust-U1
MCQ Accuracy (low)
73.53
9
1mo ago
SA-1b photos
CIN
Identity Bit Accuracy
100
9
4mo ago
Meta AI images
CIN
Identity Bit Acc
100
9
4mo ago
Perturbation Dataset
L4L
Change Accuracy
62.56
8
4mo ago
LLMBar
Qwen3-30B-A3B-Thinking-2507
Accuracy
83.07
8
4mo ago
BiasBench
Qwen2.5-32B-Instruct
Accuracy
82.5
8
4mo ago
Lexical Variation (abbr.)
Mamba
Jensen-Shannon Divergence
0.0476
8
4mo ago
CIFAR-100-C
Deep ens. (LPBN)
mCE
43.15
8
4mo ago
VizWiz
Latent Denoising
Accuracy
70.9
6
3mo ago
RWQA
Latent Denoising
Accuracy
72.9
6
3mo ago
NaturalBench
Latent Denoising
GACC
33.5
6
3mo ago
CartPole A=9.5 (test)
+DR
Average Reward
231.8
6
4mo ago
CartPole A=9.0 (test)
+ESN-OA-PT
Average Reward
830.9
6
4mo ago
CartPole A=8.5 (test)
+ESN-OA-PT
Average Reward
810.1
6
4mo ago
CartPole A=8.0 (test)
+ESN-OA-PT
Average Reward
1,000
6
4mo ago
Showing 25 of 53 rows
25 / page
50 / page
100 / page
1
2
3
Search any
task
Search any
task
Privacy Policy
Terms of Service
FAQs
Swarm Docs