Word Embedding Transformation for Robust Unsupervised Bilingual Lexicon Induction

05/26/2021
by   Hailong Cao, et al.
0

Great progress has been made in unsupervised bilingual lexicon induction (UBLI) by aligning the source and target word embeddings independently trained on monolingual corpora. The common assumption of most UBLI models is that the embedding spaces of two languages are approximately isomorphic. Therefore the performance is bound by the degree of isomorphism, especially on etymologically and typologically distant languages. To address this problem, we propose a transformation-based method to increase the isomorphism. Embeddings of two languages are made to match with each other by rotating and scaling. The method does not require any form of supervision and can be applied to any language pair. On a benchmark data set of bilingual lexicon induction, our approach can achieve competitive or superior performance compared to state-of-the-art methods, with particularly strong results being found on distant languages.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
04/04/2019

Density Matching for Bilingual Word Embedding

Recent approaches to cross-lingual word embedding have generally been ba...
research
08/19/2019

Bilingual Lexicon Induction with Semi-supervision in Non-Isometric Embedding Spaces

Recent work on bilingual lexicon induction (BLI) has frequently depended...
research
12/19/2017

Unsupervised Word Mapping Using Structural Similarities in Monolingual Embeddings

Most existing methods of automatic bilingual dictionary induction rely o...
research
10/16/2020

Multi-Adversarial Learning for Cross-Lingual Word Embeddings

Generative adversarial networks (GANs) have succeeded in inducing cross-...
research
01/31/2020

Unsupervised Bilingual Lexicon Induction Across Writing Systems

Recent embedding-based methods in unsupervised bilingual lexicon inducti...
research
11/30/2020

A Simple and Effective Approach to Robust Unsupervised Bilingual Dictionary Induction

Unsupervised Bilingual Dictionary Induction methods based on the initial...
research
10/14/2020

A Relaxed Matching Procedure for Unsupervised BLI

Recently unsupervised Bilingual Lexicon Induction (BLI) without any para...

Please sign up or login with your details

Forgot password? Click here to reset