| Task Name | Dataset Name | SOTA Result | Trend | |
|---|---|---|---|---|
| Multimodal Fact-Level Attribution | Video-MMMU 1.0 (sampled examples) | Accuracy86.8 | 24 | |
| Video Understanding | Video-MMMU (test) | Accuracy63.6 | 9 | |
| Video VQA | Video-MMMU | Score48.2 | 7 | |
| Visual Question Answering | Video-MMMU 1.0 (sampled examples) | Acc86 | 4 |