Share your thoughts, 1 month free Claude Pro on us
See more
Home
/
Benchmarks
LLM-as-a-Judge Evaluation using 10-fold Cross-Validation
Loading...
88
CG Accuracy
Qwen3-14B
22.48
39.49
56.5
73.51
Jun 4, 2026
CG Accuracy
Other Model Accuracy
CG Success Rate (Top-3)
Updated 1mo ago
Evaluation Results
Method
Method
Links
CG Accuracy
Other Model Accuracy
CG Success Rate (Top-3)
Qwen3-14B
Target=Concepts C
2026.06
88
80
82.9
Gemini-2-Flash
Target=Concepts C
2026.06
83
72
80.8
GPT-OSS 20B
Target=Prediction y^
2026.06
83
78
53
Gemini-2-Flash
Target=Prediction y^
2026.06
79
74
57.5
GPT-OSS 20B
Target=Concepts C
2026.06
77
67
86.2
Qwen3-14B
Target=Prediction y^
2026.06
64
62
64.4
Random Baseline
Target=Prediction y^
2026.06
50
50
12
Random Baseline
Target=Concepts C
2026.06
25
25
27
Feedback
Search any
task
Search any
task