| Task Name | Dataset Name | SOTA Result | Trend | |
|---|---|---|---|---|
| Multimodal Understanding and Question Answering | LLaVA-NeXT Evaluation Suite (GQA, SQA-IMG, VQA-Text, POPE, MME, MMB-EN, MMB-CN, MMVet) 13B (test) | GQA Score64.4 | 28 | |
| Multimodal Question Answering | LLaVA-NeXT-7B Evaluation Suite | FLOPs (TFLOPs)5.6 | 6 | |
| Layer Pruning Efficiency Evaluation | LLaVA-NeXT 8B | Peak VRAM (GB)18.1 | 6 | |
| Multimodal Understanding | LLaVA-NeXT Evaluation Suite | Average Accuracy100 | 6 | |
| Textual Concept Interpretability | LLaVA-NeXT textual concepts | Detection Score81 | 3 |