Share your thoughts, 1 month free Claude Pro on us
See more
Home
/
Benchmarks
LLM-as-a-judge Evaluation on Coding GPT-5.4
Loading...
67.2
Adjusted Accuracy
AURA
59.2336
61.3018
63.37
65.4382
Jun 18, 2026
Adjusted Accuracy
Performance Gain
Human Verification Count
Updated 1mo ago
Evaluation Results
Method
Method
Links
Adjusted Accuracy
Performance Gain
Human Verification Count
AURA
Judge Model=GPT-5.4, O...
2026.06
67.2
3.28
100.6
Random Forest
Judge Model=GPT-5.4, O...
2026.06
63.23
-0.69
590
MLP
Judge Model=GPT-5.4, O...
2026.06
60.83
-3.09
590
Label Prop.
Judge Model=GPT-5.4, O...
2026.06
60.5
-3.42
590
Logistic Reg.
Judge Model=GPT-5.4, O...
2026.06
59.54
-4.38
590
Feedback
Search any
task
Search any
task