OPAD: An Optimized Policy-based Active Learning Framework for Document Content Analysis

10/01/2021
by   Sumit Shekhar, et al.
4

Documents are central to many business systems, and include forms, reports, contracts, invoices or purchase orders. The information in documents is typically in natural language, but can be organized in various layouts and formats. There have been recent spurt of interest in understanding document content with novel deep learning architectures. However, document understanding tasks need dense information annotations, which are costly to scale and generalize. Several active learning techniques have been proposed to reduce the overall budget of annotation while maintaining the performance of the underlying deep learning model. However, most of these techniques work only for classification problems. But content detection is a more complex task, and has been scarcely explored in active learning literature. In this paper, we propose OPAD, a novel framework using reinforcement policy for active learning in content detection tasks for documents. The proposed framework learns the acquisition function to decide the samples to be selected while optimizing performance metrics that the tasks typically have. Furthermore, we extend to weak labelling scenarios to further reduce the cost of annotation significantly. We propose novel rewards to account for class imbalance and user feedback in the annotation interface, to improve the active learning method. We show superior performance of the proposed OPAD framework for active learning for various tasks related to document understanding like layout parsing, object detection and named entity recognition. Ablation studies for human feedback and class imbalance rewards are presented, along with a comparison of annotation times for different approaches.

READ FULL TEXT
research
11/02/2022

Improving Named Entity Recognition in Telephone Conversations via Effective Active Learning with Human in the Loop

Telephone transcription data can be very noisy due to speech recognition...
research
01/08/2020

LTP: A New Active Learning Strategy for Bert-CRF Based Named Entity Recognition

In recent years, deep learning has achieved great success in many natura...
research
04/18/2022

Active Learning with Weak Labels for Gaussian Processes

Annotating data for supervised learning can be costly. When the annotati...
research
08/07/2019

An Adaptive Supervision Framework for Active Learning in Object Detection

Active learning approaches in computer vision generally involve querying...
research
08/20/2021

Region-level Active Learning for Cluttered Scenes

Active learning for object detection is conventionally achieved by apply...
research
06/26/2018

A Practical Incremental Learning Framework For Sparse Entity Extraction

This work addresses challenges arising from extracting entities from tex...
research
10/02/2021

Automated Seed Quality Testing System using GAN Active Learning

Quality assessment of agricultural produce is a crucial step in minimizi...

Please sign up or login with your details

Forgot password? Click here to reset