Exploiting Lists of Names for Named Entity Identification of Financial Institutions from Unstructured Documents

by   Zheng Xu, et al.

There is a wealth of information about financial systems that is embedded in document collections. In this paper, we focus on a specialized text extraction task for this domain. The objective is to extract mentions of names of financial institutions, or FI names, from financial prospectus documents, and to identify the corresponding real world entities, e.g., by matching against a corpus of such entities. The tasks are Named Entity Recognition (NER) and Entity Resolution (ER); both are well studied in the literature. Our contribution is to develop a rule-based approach that will exploit lists of FI names for both tasks; our solution is labeled Dict-based NER and Rank-based ER. Since the FI names are typically represented by a root, and a suffix that modifies the root, we use these lists of FI names to create specialized root and suffix dictionaries. To evaluate the effectiveness of our specialized solution for extracting FI names, we compare Dict-based NER with a general purpose rule-based NER solution, ORG NER. Our evaluation highlights the benefits and limitations of specialized versus general purpose approaches, and presents additional suggestions for tuning and customization for FI name extraction. To our knowledge, our proposed solutions, Dict-based NER and Rank-based ER, and the root and suffix dictionaries, are the first attempt to exploit specialized knowledge, i.e., lists of FI names, for rule-based NER and ER.


page 1

page 2

page 3

page 4


Man is to Person as Woman is to Location: Measuring Gender Bias in Named Entity Recognition

We study the bias in several state-of-the-art named entity recognition (...

Named Entity Recognition in the Legal Domain using a Pointer Generator Network

Named Entity Recognition (NER) is the task of identifying and classifyin...

Automated Testing and Improvement of Named Entity Recognition Systems

Named entity recognition (NER) systems have seen rapid progress in recen...

Biomedical Named Entity Recognition via Reference-Set Augmented Bootstrapping

We present a weakly-supervised data augmentation approach to improve Nam...

Old Content and Modern Tools - Searching Named Entities in a Finnish OCRed Historical Newspaper Collection 1771-1910

Named Entity Recognition (NER), search, classification and tagging of na...

Extracting Impact Model Narratives from Social Services' Text

Named entity recognition (NER) is an important task in narration extract...

EasyNER: A Customizable Easy-to-Use Pipeline for Deep Learning- and Dictionary-based Named Entity Recognition from Medical Text

Medical research generates a large number of publications with the PubMe...

Please sign up or login with your details

Forgot password? Click here to reset