Disambiguation of Company names via Deep Recurrent Networks

03/07/2023
by   Alessandro Basile, et al.
0

Name Entity Disambiguation is the Natural Language Processing task of identifying textual records corresponding to the same Named Entity, i.e. real-world entities represented as a list of attributes (names, places, organisations, etc.). In this work, we face the task of disambiguating companies on the basis of their written names. We propose a Siamese LSTM Network approach to extract – via supervised learning – an embedding of company name strings in a (relatively) low dimensional vector space and use this representation to identify pairs of company names that actually represent the same company (i.e. the same Entity). Given that the manual labelling of string pairs is a rather onerous task, we analyse how an Active Learning approach to prioritise the samples to be labelled leads to a more efficient overall learning pipeline. With empirical investigations, we show that our proposed Siamese Network outperforms several benchmark approaches based on standard string matching algorithms when enough labelled data are available. Moreover, we show that Active Learning prioritisation is indeed helpful when labelling resources are limited, and let the learning models reach the out-of-sample performance saturation with less labelled data with respect to standard (random) data labelling approaches.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
10/30/2020

Learning Structured Representations of Entity Names using Active Learning and Weak Supervision

Structured representations of entity names are useful for many entity-re...
research
09/05/2018

Merging datasets through deep learning

Merging datasets is a key operation for data analytics. A frequent requi...
research
07/19/2019

Fast Record Linkage for Company Entities

Record Linkage is an essential part of almost all real-world systems tha...
research
07/11/2023

Named entity recognition using GPT for identifying comparable companies

For both public and private firms, comparable companies analysis is wide...
research
04/30/2020

Named Entity Recognition without Labelled Data: A Weak Supervision Approach

Named Entity Recognition (NER) performance often degrades rapidly when a...
research
01/27/2021

On Statistical Bias In Active Learning: How and When To Fix It

Active learning is a powerful tool when labelling data is expensive, but...
research
07/22/2017

Identifying civilians killed by police with distantly supervised entity-event extraction

We propose a new, socially-impactful task for natural language processin...

Please sign up or login with your details

Forgot password? Click here to reset