End-to-End Information Extraction by Character-Level Embedding and Multi-Stage Attentional U-Net

06/02/2021
by   Tuan-Anh Nguyen Dang, et al.
0

Information extraction from document images has received a lot of attention recently, due to the need for digitizing a large volume of unstructured documents such as invoices, receipts, bank transfers, etc. In this paper, we propose a novel deep learning architecture for end-to-end information extraction on the 2D character-grid embedding of the document, namely the Multi-Stage Attentional U-Net. To effectively capture the textual and spatial relations between 2D elements, our model leverages a specialized multi-stage encoder-decoders design, in conjunction with efficient uses of the self-attention mechanism and the box convolution. Experimental results on different datasets show that our model outperforms the baseline U-Net architecture by a large margin while using 40% fewer parameters. Moreover, it also significantly improved the baseline in erroneous OCR and limited training data scenario, thus becomes practical for real-world applications.

READ FULL TEXT

page 1

page 4

page 10

research
07/11/2022

GMN: Generative Multi-modal Network for Practical Document Information Extraction

Document Information Extraction (DIE) has attracted increasing attention...
research
09/12/2020

Abstractive Information Extraction from Scanned Invoices (AIESI) using End-to-end Sequential Approach

Recent proliferation in the field of Machine Learning and Deep Learning ...
research
12/06/2019

NASNet: A Neuron Attention Stage-by-Stage Net for Single Image Deraining

Images captured under complicated rain conditions often suffer from noti...
research
07/13/2022

Imaging through the Atmosphere using Turbulence Mitigation Transformer

Restoring images distorted by atmospheric turbulence is a long-standing ...
research
07/03/2019

AMI-Net+: A Novel Multi-Instance Neural Network for Medical Diagnosis from Incomplete and Imbalanced Data

In medical real-world study (RWS), how to fully utilize the fragmentary ...
research
08/23/2021

Using Neighborhood Context to Improve Information Extraction from Visual Documents Captured on Mobile Phones

Information Extraction from visual documents enables convenient and inte...
research
11/14/2019

Character Keypoint-based Homography Estimation in Scanned Documents for Efficient Information Extraction

Precise homography estimation between multiple images is a pre-requisite...

Please sign up or login with your details

Forgot password? Click here to reset