| Task Name | Dataset Name | SOTA Result | Trend | |
|---|---|---|---|---|
| Visual Question Answering | Multimodal Benchmark Suite Aggregate | Average Ratio100 | 13 | |
| Multimodal Question Answering and Understanding | Multimodal Benchmark Suite (MMBench-EN, MMVet, TextVQA, SQA-IMG, GQA, VQAv2, SEED-Image, POPE) | MMBench-EN68.5 | 11 |