Share your thoughts, 1 month free Claude Pro on us
See more
Home
/
Benchmarks
Multimodal Reasoning on SAT
Loading...
59.49
Accuracy
Ocean_R1_3B_Instruct
22.7052
32.2551
41.805
51.3549
Oct 24, 2025
Accuracy
Updated 1mo ago
Evaluation Results
Method
Method
Links
Accuracy
Ocean_R1_3B_Instruct
2025.10
59.49
RECAP
backbone=Qwen2.5-VL-3B
2025.10
55.19
MM-R1-MGT-PerceReason
2025.10
50.83
vision-grpo-qwen-2.5-vl-3b
2025.10
50.57
Qwen2.5-VL-3B
variant=MoDoMoDo
2025.10
49.95
VLAA-Thinker-3B
2025.10
49.38
Qwen2.5-VL-3B
variant=Uniform
2025.10
44.55
Qwen2.5-VL-3B
mode=Base model
2025.10
43.98
Qwen2.5-VL-3B-Instruct-GRPO-deepmath
2025.10
34.7
Qwen2.5VL-3b-RLCS
2025.10
24.12
Feedback
Search any
task
Search any
task