ERNIE-Layout: Layout Knowledge Enhanced Pre-training for Visually-rich Document Understanding

About

Recent years have witnessed the rise and success of pre-training techniques in visually-rich document understanding. However, most existing methods lack the systematic mining and utilization of layout-centered knowledge, leading to sub-optimal performances. In this paper, we propose ERNIE-Layout, a novel document pre-training solution with layout knowledge enhancement in the whole workflow, to learn better representations that combine the features from text, layout, and image. Specifically, we first rearrange input sequences in the serialization stage, and then present a correlative pre-training task, reading order prediction, to learn the proper reading order of documents. To improve the layout awareness of the model, we integrate a spatial-aware disentangled attention into the multi-modal transformer and a replaced regions prediction task into the pre-training phase. Experimental results show that ERNIE-Layout achieves superior performance on various downstream tasks, setting new state-of-the-art on key information extraction, document image classification, and document question answering datasets. The code and models are publicly available at http://github.com/PaddlePaddle/PaddleNLP/tree/develop/model_zoo/ernie-layout.

Qiming Peng, Yinxu Pan, Wenjin Wang, Bin Luo, Zhenyu Zhang, Zhengjie Huang, Teng Hu, Weichong Yin, Yongfeng Chen, Yin Zhang, Shikun Feng, Yu Sun, Hao Tian, Hua Wu, Haifeng Wang• 2022

Related benchmarks

Task	Dataset	Result
Document Classification	RVL-CDIP (test)	Accuracy96.27	306
Document Visual Question Answering	DocVQA (test)	ANLS84.86	292
Information Extraction	CORD (test)	F1 Score97.21	133
Entity extraction	FUNSD (test)	Entity F1 Score93.12	104
Form Understanding	FUNSD (test)	F1 Score90.28	73
Document Question Answering	DocVQA	ANLS88.41	64
Information Extraction	SROIE (test)	F1 Score97.55	62
Visual Question Answering	DocVQA	ANLS88.4	59
Semantic Entity Recognition	CORD	F1 Score97.21	55
Semantic Entity Recognition	FUNSD	EN Score93.12	31

Showing 10 of 14 rows

Other info

Code

Follow for update

@wizwand_team Discord