Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

LLaVA-NeXT

Benchmarks

Task NameDataset NameSOTA ResultTrend
Multimodal Understanding and Question AnsweringLLaVA-NeXT Evaluation Suite (GQA, SQA-IMG, VQA-Text, POPE, MME, MMB-EN, MMB-CN, MMVet) 13B (test)
GQA Score64.4
28
Multimodal Question AnsweringLLaVA-NeXT-7B Evaluation Suite
FLOPs (TFLOPs)5.6
6
Layer Pruning Efficiency EvaluationLLaVA-NeXT 8B
Peak VRAM (GB)18.1
6
Multimodal UnderstandingLLaVA-NeXT Evaluation Suite
Average Accuracy100
6
Textual Concept InterpretabilityLLaVA-NeXT textual concepts
Detection Score81
3
Showing 5 of 5 rows