Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

MMLongBench

Benchmarks

Task NameDataset NameSOTA ResultTrend
Multimodal Document Question AnsweringMMLongBench-Doc
Overall Accuracy65.8
77
Long-context document understandingMMLongBench-Doc
Accuracy55.8
58
Document Visual Question AnsweringMMLongbench doc
Accuracy45.6
48
Visual Document RetrievalMMLongBench
Doc Retrieval Rate53.82
46
Document Question AnsweringMMLongBench-Doc
Accuracy65.8
40
Multimodal Document Question AnsweringMMLongBench
Accuracy48.2
26
Long-document Visual Question AnsweringMMLongBench Overall
Average Score90.77
22
Long-document Visual Question AnsweringMMLongBench 128K context
MMLB-D83.33
22
Long-document Visual Question AnsweringMMLongBench 64K context
MMLB-D93.1
22
RetrievalMMLongBench
Recall75.86
18
Long-context Multi-modal UnderstandingMMLongBench
Text Accuracy27.49
17
Document Question AnsweringMMLongBench-Doc (test)
Accuracy49.09
16
Document Understanding, OCR & ChartsMMLongBench Doc
Score57.5
14
Evidence-page retrievalMMLongBench-Doc
Recall75.68
12
Reasoning over rich modalitiesMMLongBench Doc
Accuracy42.3
12
Multimodal Document Question AnsweringMMLongBench (test)
Chart Acc.34.7
12
Long-context Visual Question AnsweringMMLongBench 32K
Accuracy82.4
11
Long-context Visual Question AnsweringMMLongBench 128K
Accuracy78.6
11
Document Question AnsweringMMLongBench
Exact Match43.8
11
RetrievalMMLongBench Finreport
MRR@1049.62
6
RetrievalMMLongBench Doc
MRR@1047.64
6
LongContext UnderstandingMMLongBench-Doc
Pass@161.4
5
LongContext UnderstandingMMLongBench
Pass@174.8
5
Dataset Description ExtractionMMLongBench-Doc
Accuracy94.9
5
Long-document Visual Question AnsweringMMLongBench 512K context
MMLongBench-D Score31.91
4
Showing 25 of 39 rows