A Survey of Code-switching: Linguistic and Social Perspectives for Language Technologies

01/05/2023
by   A. Seza Doğruöz, et al.
0

The analysis of data in which multiple languages are represented has gained popularity among computational linguists in recent years. So far, much of this research focuses mainly on the improvement of computational methods and largely ignores linguistic and social aspects of C-S discussed across a wide range of languages within the long-established literature in linguistics. To fill this gap, we offer a survey of code-switching (C-S) covering the literature in linguistics with a reflection on the key issues in language technologies. From the linguistic perspective, we provide an overview of structural and functional patterns of C-S focusing on the literature from European and Indian contexts as highly multilingual areas. From the language technologies perspective, we discuss how massive language models fail to represent diverse C-S types due to lack of appropriate training data, lack of robust evaluation benchmarks for C-S (across multilingual situations and types of C-S) and lack of end-to-end systems that cover sociolinguistic aspects of C-S as well. Our survey will be a step towards an outcome of mutual benefit for computational scientists and linguists with a shared interest in multilingualism and C-S.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
05/09/2020

LinCE: A Centralized Benchmark for Linguistic Code-switching Evaluation

Recent trends in NLP research have raised an interest in linguistic code...
research
03/25/2019

A Survey of Code-switched Speech and Language Processing

Code-switching, the alternation of languages within a conversation or ut...
research
10/11/2016

Survey on the Use of Typological Information in Natural Language Processing

In recent years linguistic typology, which classifies the world's langua...
research
02/28/2021

RuSentEval: Linguistic Source, Encoder Force!

The success of pre-trained transformer language models has brought a gre...
research
07/01/2021

A Primer on Pretrained Multilingual Language Models

Multilingual Language Models (MLLMs) such as mBERT, XLM, XLM-R, etc. hav...
research
07/25/2023

Diversity and Language Technology: How Techno-Linguistic Bias Can Cause Epistemic Injustice

It is well known that AI-based language technology – large language mode...
research
05/23/2023

TalkUp: A Novel Dataset Paving the Way for Understanding Empowering Language

Empowering language is important in many real-world contexts, from educa...

Please sign up or login with your details

Forgot password? Click here to reset