Share your thoughts, 1 month free Claude Pro on us
See more
Home
/
Benchmarks
Multimodal Question Answering on Frontier-comparable pool 20-eval macro mean (test)
Loading...
10.23
AME Change (pp)
InternVL3-2B
2.0764
4.1932
6.31
8.4268
Jun 24, 2026
AME Change (pp)
Datology FLOPs/correct
Comparator FLOPs/correct
Cost ratio
Updated 1mo ago
Evaluation Results
Method
Method
Links
AME Change (pp)
Datology FLOPs/correct
Comparator FLOPs/correct
Cost ratio
InternVL3-2B
Length overlap=86%
2026.06
10.23
0.27
0.21
0.8
InternVL3.5-2B
Length overlap=26%
2026.06
9.68
0.27
1.05
3.9
Qwen3.5-2B
Length overlap=48%
2026.06
6.99
0.27
0.97
3.6
InternVL3.5-4B
Length overlap=53%
2026.06
5.69
0.41
1.12
2.7
Qwen3-VL-4B-Instruct
Length overlap=81%
2026.06
4.53
0.41
1.3
3.2
Qwen3-VL-2B-Instruct
Length overlap=89%
2026.06
3.93
0.27
0.63
2.4
Perceptron-Isaac-2B
Length overlap=18%
2026.06
2.39
0.27
0.91
2.2
Feedback
Search any
task
Search any
task