| Task Name | Dataset Name | SOTA Result | Trend | |
|---|---|---|---|---|
| Multimodal Understanding | MMMU | Accuracy81.8 | 437 | |
| Multi-discipline Multimodal Understanding | MMMU | Accuracy84.2 | 422 | |
| Massive Multi-discipline Multimodal Understanding | MMMU | Accuracy71.5 | 249 | |
| Multimodal Understanding | MMMU | MMMU Score67.8 | 232 | |
| Multimodal Reasoning | MMMU | Accuracy83.89 | 220 | |
| Multi-discipline Multimodal Understanding | MMMU (val) | Accuracy81.7 | 212 | |
| Multimodal Understanding | MMMU (val) | MMMU Score85.2 | 211 | |
| Multimodal Reasoning | MMMU Pro | Accuracy85.6 | 171 | |
| Multimodal Reasoning | MMMU (val) | Accuracy78.2 | 168 | |
| Multimodal Understanding | MMMU (test) | MMMU Score69.6 | 112 | |
| Multimodal Understanding | MMMU | MMMU Score81.8 | 110 | |
| Multimodal Understanding | MMMU | Accuracy59.63 | 107 | |
| Visual Question Answering | MMMU | Accuracy81.7 | 101 | |
| Multi-modal Question Answering | MMMU | Accuracy82.3 | 98 | |
| Multi-discipline Reasoning | MMMU | Accuracy36.1 | 83 | |
| Video reasoning | Video-MMMU | Accuracy84.6 | 83 | |
| Multimodal Reasoning | MMMU | Accuracy72.9 | 77 | |
| Multimodal Understanding | MMMU | Accuracy (MMMU)83.4 | 73 | |
| Vision Understanding | MMMU | Accuracy72.9 | 71 | |
| Multimodal Understanding | MMMU | MMMU Score60.74 | 69 | |
| Multi-discipline Multimodal Understanding | MMMU Pro | Accuracy67.3 | 66 | |
| General Reasoning | MMMU | Overall Score85.4 | 57 | |
| Multi-agent discussion attack | MMMU | Delta Accuracy2.3 | 48 | |
| Multimodal Reasoning | MMMU (test) | Accuracy64.7 | 43 | |
| Multimodal Reasoning | MMMU | Accuracy85.79 | 40 |