Exploiting Lists of Names for Named Entity Identification of Financial Institutions from Unstructured Documents

02/14/2016
by   Zheng Xu, et al.
0

There is a wealth of information about financial systems that is embedded in document collections. In this paper, we focus on a specialized text extraction task for this domain. The objective is to extract mentions of names of financial institutions, or FI names, from financial prospectus documents, and to identify the corresponding real world entities, e.g., by matching against a corpus of such entities. The tasks are Named Entity Recognition (NER) and Entity Resolution (ER); both are well studied in the literature. Our contribution is to develop a rule-based approach that will exploit lists of FI names for both tasks; our solution is labeled Dict-based NER and Rank-based ER. Since the FI names are typically represented by a root, and a suffix that modifies the root, we use these lists of FI names to create specialized root and suffix dictionaries. To evaluate the effectiveness of our specialized solution for extracting FI names, we compare Dict-based NER with a general purpose rule-based NER solution, ORG NER. Our evaluation highlights the benefits and limitations of specialized versus general purpose approaches, and presents additional suggestions for tuning and customization for FI name extraction. To our knowledge, our proposed solutions, Dict-based NER and Rank-based ER, and the root and suffix dictionaries, are the first attempt to exploit specialized knowledge, i.e., lists of FI names, for rule-based NER and ER.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
10/24/2019

Man is to Person as Woman is to Location: Measuring Gender Bias in Named Entity Recognition

We study the bias in several state-of-the-art named entity recognition (...
research
12/17/2020

Named Entity Recognition in the Legal Domain using a Pointer Generator Network

Named Entity Recognition (NER) is the task of identifying and classifyin...
research
08/14/2023

Automated Testing and Improvement of Named Entity Recognition Systems

Named entity recognition (NER) systems have seen rapid progress in recen...
research
06/01/2019

Biomedical Named Entity Recognition via Reference-Set Augmented Bootstrapping

We present a weakly-supervised data augmentation approach to improve Nam...
research
11/09/2016

Old Content and Modern Tools - Searching Named Entities in a Finnish OCRed Historical Newspaper Collection 1771-1910

Named Entity Recognition (NER), search, classification and tagging of na...
research
04/04/2022

Extracting Impact Model Narratives from Social Services' Text

Named entity recognition (NER) is an important task in narration extract...
research
04/16/2023

EasyNER: A Customizable Easy-to-Use Pipeline for Deep Learning- and Dictionary-based Named Entity Recognition from Medical Text

Medical research generates a large number of publications with the PubMe...

Please sign up or login with your details

Forgot password? Click here to reset