Share your thoughts, 1 month free Claude Pro on us
See more
Home
/
Benchmarks
Self-correction on Llama-3.3-70B math n=30 (failure pool)
Loading...
86.7
Correction Rate
L_MEMORY (Source-Conditioned Role Relabeling)
-3.468
19.941
43.35
66.759
Jun 4, 2026
Correction Rate
Updated 1mo ago
Evaluation Results
Method
Method
Links
Correction Rate
L_MEMORY (Source-Conditioned Role Relabeling)
Protocol=L_MEM.
2026.06
86.7
Reflexion
Protocol=Reflex.
2026.06
6.7
audit
Protocol=audit-only ba...
2026.06
3.3
Self-Refine
Protocol=S.-Refine
2026.06
0
Chain-of-Verification (CoVe)
Protocol=CoVe
2026.06
0
Feedback
Search any
task
Search any
task