XLEnt: Mining a Large Cross-lingual Entity Dataset with Lexical-Semantic-Phonetic Word Alignment

04/17/2021
by   Ahmed El-Kishky, et al.
0

Cross-lingual named-entity lexicon are an important resource to multilingual NLP tasks such as machine translation and cross-lingual wikification. While knowledge bases contain a large number of entities in high-resource languages such as English and French, corresponding entities for lower-resource languages are often missing. To address this, we propose Lexical-Semantic-Phonetic Align (LSP-Align), a technique to automatically mine cross-lingual entity lexicon from the web. We demonstrate LSP-Align outperforms baselines at extracting cross-lingual entity pairs and mine 164 million entity pairs from 120 different languages aligned with English. We release these cross-lingual entity pairs along with the massively multilingual tagged named entity corpus as a resource to the NLP community.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
12/17/2019

Cross-Lingual Ability of Multilingual BERT: An Empirical Study

Recent work has exhibited the surprising cross-lingual abilities of mult...
research
06/17/2022

Statistical and Neural Methods for Cross-lingual Entity Label Mapping in Knowledge Graphs

Knowledge bases such as Wikidata amass vast amounts of named entity info...
research
06/18/2018

Co-training Embeddings of Knowledge Graphs and Entity Descriptions for Cross-lingual Entity Alignment

Multilingual knowledge graph (KG) embeddings provide latent semantic rep...
research
11/27/2018

Joint Representation Learning of Cross-lingual Words and Entities via Attentive Distant Supervision

Joint representation learning of words and entities benefits many NLP ta...
research
10/15/2021

Cross-Lingual Fine-Grained Entity Typing

The growth of cross-lingual pre-trained models has enabled NLP tools to ...
research
08/14/2016

Proceedings of the LexSem+Logics Workshop 2016

Lexical semantics continues to play an important role in driving researc...
research
08/31/2019

Adversarial Learning with Contextual Embeddings for Zero-resource Cross-lingual Classification and NER

Contextual word embeddings (e.g. GPT, BERT, ELMo, etc.) have demonstrate...

Please sign up or login with your details

Forgot password? Click here to reset