Information Extraction in Domain and Generic Documents: Findings from Heuristic-based and Data-driven Approaches

06/30/2023
by   Shiyu Yuan, et al.
0

Information extraction (IE) plays very important role in natural language processing (NLP) and is fundamental to many NLP applications that used to extract structured information from unstructured text data. Heuristic-based searching and data-driven learning are two main stream implementation approaches. However, no much attention has been paid to document genre and length influence on IE tasks. To fill the gap, in this study, we investigated the accuracy and generalization abilities of heuristic-based searching and data-driven to perform two IE tasks: named entity recognition (NER) and semantic role labeling (SRL) on domain-specific and generic documents with different length. We posited two hypotheses: first, short documents may yield better accuracy results compared to long documents; second, generic documents may exhibit superior extraction outcomes relative to domain-dependent documents due to training document genre limitations. Our findings reveals that no single method demonstrated overwhelming performance in both tasks. For named entity extraction, data-driven approaches outperformed symbolic methods in terms of accuracy, particularly in short texts. In the case of semantic roles extraction, we observed that heuristic-based searching method and data-driven based model with syntax representation surpassed the performance of pure data-driven approach which only consider semantic information. Additionally, we discovered that different semantic roles exhibited varying accuracy levels with the same method. This study offers valuable insights for downstream text mining tasks, such as NER and SRL, when addressing various document features and genres.

READ FULL TEXT

page 1

page 7

page 8

research
08/31/2021

TNNT: The Named Entity Recognition Toolkit

Extraction of categorised named entities from text is a complex task giv...
research
03/04/2020

Kleister: A novel task for Information Extraction involving Long Documents with Complex Layout

State-of-the-art solutions for Natural Language Processing (NLP) are abl...
research
10/19/2018

Efficient Dependency-Guided Named Entity Recognition

Named entity recognition (NER), which focuses on the extraction of seman...
research
08/06/2021

Lights, Camera, Action! A Framework to Improve NLP Accuracy over OCR documents

Document digitization is essential for the digital transformation of our...
research
07/18/2023

Mutual Reinforcement Effects in Japanese Sentence Classification and Named Entity Recognition Tasks

Information extraction(IE) is a crucial subfield within natural language...
research
11/06/2021

Focusing on Possible Named Entities in Active Named Entity Label Acquisition

Named entity recognition (NER) aims to identify mentions of named entiti...
research
03/29/2019

CUTIE: Learning to Understand Documents with Convolutional Universal Text Information Extractor

Extracting key information from documents, such as receipts or invoices,...

Please sign up or login with your details

Forgot password? Click here to reset