Share your thoughts, 1 month free Claude Pro on us
See more
Home
/
Benchmarks
Veracity Prediction on Factiverse Multilingual Evaluation Set Aggregate Mean (test)
Loading...
64.57
Macro F1
GPT-5.2
59.9836
61.1743
62.365
63.5557
Jun 7, 2026
Macro F1
Micro F1
Updated 1mo ago
Evaluation Results
Method
Method
Links
Macro F1
Micro F1
GPT-5.2
2026.06
64.57
66.55
Factiverse
2026.06
62.05
64.57
Claude Opus 4.6
2026.06
61.22
64.17
Qwen3-8b
2026.06
60.16
60.92
Feedback
Search any
task
Search any
task