Crosslingual Transfer Learning for Low-Resource Languages Based on Multilingual Colexification Graphs

05/22/2023
by   Yihong Liu, et al.
0

Colexification in comparative linguistics refers to the phenomenon of a lexical form conveying two or more distinct meanings. In this paper, we propose simple and effective methods to build multilingual graphs from colexification patterns: ColexNet and ColexNet+. ColexNet's nodes are concepts and its edges are colexifications. In ColexNet+, concept nodes are in addition linked through intermediate nodes, each representing an ngram in one of 1,334 languages. We use ColexNet+ to train high-quality multilingual embeddings that are well-suited for transfer learning scenarios. Existing work on colexification patterns relies on annotated word lists. This limits scalability and usefulness in NLP. In contrast, we identify colexification patterns of more than 2,000 concepts across 1,335 languages directly from an unannotated parallel corpus. In our experiments, we first show that ColexNet has a high recall on CLICS, a dataset of crosslingual colexifications. We then evaluate on roundtrip translation, verse retrieval and verse classification and show that our embeddings surpass several baselines in a transfer learning setting. This demonstrates the benefits of colexification for multilingual NLP.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
05/27/2023

Enhancing Translation for Indigenous Languages: Experiments with Multilingual Models

This paper describes CIC NLP's submission to the AmericasNLP 2023 Shared...
research
12/15/2019

A Comparison of Architectures and Pretraining Methods for Contextualized Multilingual Word Embeddings

The lack of annotated data in many languages is a well-known challenge w...
research
10/07/2020

Transfer Learning and Distant Supervision for Multilingual Transformer Models: A Study on African Languages

Multilingual transformer models like mBERT and XLM-RoBERTa have obtained...
research
07/14/2021

ParCourE: A Parallel Corpus Explorer for a Massively Multilingual Corpus

With more than 7000 languages worldwide, multilingual natural language p...
research
11/08/2019

Instance-based Transfer Learning for Multilingual Deep Retrieval

Perhaps the simplest type of multilingual transfer learning is instance-...
research
05/23/2023

Beyond Shared Vocabulary: Increasing Representational Word Similarities across Languages for Multilingual Machine Translation

Using a shared vocabulary is common practice in Multilingual Neural Mach...
research
12/04/2019

Towards Building a Multilingual Sememe Knowledge Base: Predicting Sememes for BabelNet Synsets

A sememe is defined as the minimum semantic unit of human languages. Sem...

Please sign up or login with your details

Forgot password? Click here to reset