Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Hierarchical Visual Feature Aggregation for OCR-Free Document Understanding

About

We present a novel OCR-free document understanding framework based on pretrained Multimodal Large Language Models (MLLMs). Our approach employs multi-scale visual features to effectively handle various font sizes within document images. To address the increasing costs of considering the multi-scale visual inputs for MLLMs, we propose the Hierarchical Visual Feature Aggregation (HVFA) module, designed to reduce the number of input tokens to LLMs. Leveraging a feature pyramid with cross-attentive pooling, our approach effectively manages the trade-off between information loss and efficiency without being affected by varying document image sizes. Furthermore, we introduce a novel instruction tuning task, which facilitates the model's text-reading capability by learning to predict the relative positions of input text, eventually minimizing the risk of truncated text caused by the limited capacity of LLMs. Comprehensive experiments validate the effectiveness of our approach, demonstrating superior performance in various document understanding tasks.

Jaeyoo Park, Jin Young Choi, Jeonghyung Park, Bohyung Han• 2024

Related benchmarks

TaskDatasetResultRank
Text-based Visual Question AnsweringTextVQA (val)
Accuracy59.2
262
Document Visual Question AnsweringDocVQA (test)
ANLS72.7
213
Chart Question AnsweringChartQA (test)
Accuracy63.3
176
Table Fact VerificationTabFact (test)
Accuracy68.2
136
Information Visual Question AnsweringInfoVQA (test)
ANLS45.9
130
Visual Question AnsweringTextVQA (test)
Accuracy59.2
124
Table Question AnsweringWikiTableQuestions (test)
Accuracy34.5
86
Table Question AnsweringWTQ (test)
Denotation Accuracy34.5
62
Image CaptioningTextCaps (test)
CIDEr135.2
50
Document Visual Question AnsweringDocVQA v1.0 (test)
ANLS72.7
49
Showing 10 of 15 rows

Other info

Follow for update