Share your thoughts, 1 month free Claude Pro on us
See more
Home
/
Benchmarks
Literary Translation on par3 dev evaluation (full)
Loading...
0.773
COMETKiwi
Gemini
0.72412
0.73681
0.7495
0.76219
Jun 24, 2026
COMETKiwi
COMET-22
MetricX
MetricX-QE
Updated 1mo ago
Evaluation Results
Method
Method
Links
COMETKiwi
COMET-22
MetricX
MetricX-QE
Gemini
Pipeline=P2, Human eva...
2026.06
0.773
0.805
5.182
4.758
Gemini
Pipeline=P1, Human eva...
2026.06
0.772
0.804
5.243
4.813
GPT-5.4
Pipeline=P2, Human eva...
2026.06
0.768
0.8
5.482
5.02
GPT-5.4
Pipeline=P1, Human eva...
2026.06
0.766
0.798
5.542
5.073
Agents
Pipeline=P3, Human eva...
2026.06
0.766
0.793
5.656
5.109
HT
Human eval.=true
2026.06
0.726
-
-
5.726
Feedback
Search any
task
Search any
task