Share your thoughts, 1 month free Claude Pro on us
See more
Home
/
Benchmarks
LLM-as-a-Judge Evaluation on Coding Qwen (n=8884)
Loading...
77.74
Adjusted Accuracy
AURA
52.8112
59.2831
65.755
72.2269
Jun 18, 2026
Adjusted Accuracy
Gain
Human Verification Count
Updated 1mo ago
Evaluation Results
Method
Method
Links
Adjusted Accuracy
Gain
Human Verification Count
AURA
Judge Model=Qwen, Orig...
2026.06
77.74
24.02
278.8
Random Forest
Judge Model=Qwen, Orig...
2026.06
59.2
5.47
1,777
Logistic Reg.
Judge Model=Qwen, Orig...
2026.06
58.47
4.75
1,777
MLP
Judge Model=Qwen, Orig...
2026.06
57.44
3.72
1,777
Label Prop.
Judge Model=Qwen, Orig...
2026.06
53.77
0.04
1,777
Feedback
Search any
task
Search any
task