Our new X account is live! Follow @wizwand_team for updates
Home
/
Benchmarks
Natural Language Inference on MultiMed-X EN
Loading...
78.67
Accuracy
GPT-4o
68.9564
71.4782
74
76.5218
Jan 13, 2026
Accuracy
Updated 4d ago
Evaluation Results
Method
Method
Links
Accuracy
GPT-4o
2026.01
78.67
MED-COREASONER
backbone=GPT-5.1
2026.01
77.33
GPT-5.2
2026.01
76.67
GPT-5.1
2026.01
76.67
Claude-3.5-haiku
2026.01
69.33
Feedback
Search any
task
Search any
task