Our new X account is live! Follow @wizwand_team for updates
WorkDL logo mark

Cost-effective End-to-end Information Extraction for Semi-structured Document Images

About

A real-world information extraction (IE) system for semi-structured document images often involves a long pipeline of multiple modules, whose complexity dramatically increases its development and maintenance cost. One can instead consider an end-to-end model that directly maps the input to the target output and simplify the entire process. However, such generation approach is known to lead to unstable performance if not designed carefully. Here we present our recent effort on transitioning from our existing pipeline-based IE system to an end-to-end system focusing on practical challenges that are associated with replacing and deploying the system in real, large-scale production. By carefully formulating document IE as a sequence generation task, we show that a single end-to-end IE system can be built and still achieve competent performance.

Wonseok Hwang, Hyunji Lee, Jinyeong Yim, Geewook Kim, Minjoon Seo• 2021

Related benchmarks

TaskDatasetResultRank
Document Information ExtractionCORD 45 (test)
F1 Score43.3
7
Document Information ExtractionReceipt
F1 Score71.5
5
Document Information ExtractionTicket 12 (test)
F1 Score41.8
5
Document Information ExtractionBusiness Card
F1 Score29.9
5
Showing 4 of 4 rows

Other info

Follow for update