A Survey of Word Reordering in Statistical Machine Translation: Computational Models and Language Phenomena

02/17/2015
by   Arianna Bisazza, et al.
0

Word reordering is one of the most difficult aspects of statistical machine translation (SMT), and an important factor of its quality and efficiency. Despite the vast amount of research published to date, the interest of the community in this problem has not decreased, and no single method appears to be strongly dominant across language pairs. Instead, the choice of the optimal approach for a new translation task still seems to be mostly driven by empirical trials. To orientate the reader in this vast and complex research area, we present a comprehensive survey of word reordering viewed as a statistical modeling challenge and as a natural language phenomenon. The survey describes in detail how word reordering is modeled within different string-based and tree-based SMT frameworks and as a stand-alone task, including systematic overviews of the literature in advanced reordering modeling. We then question why some approaches are more successful than others in different language pairs. We argue that, besides measuring the amount of reordering, it is important to understand which kinds of reordering occur in a given language pair. To this end, we conduct a qualitative analysis of word reordering phenomena in a diverse sample of language pairs, based on a large collection of linguistic knowledge. Empirical results in the SMT literature are shown to support the hypothesis that a few linguistic facts can be very useful to anticipate the reordering characteristics of a language pair and to select the SMT framework that best suits them.

READ FULL TEXT
research
10/24/2016

Statistical Machine Translation for Indian Languages: Mission Hindi

This paper discusses Centre for Development of Advanced Computing Mumbai...
research
01/16/2017

Machine Translation Approaches and Survey for Indian Languages

In this study, we present an analysis regarding the performance of the s...
research
12/16/2016

Neural Networks Classifier for Data Selection in Statistical Machine Translation

We address the data selection problem in statistical machine translation...
research
08/16/2016

Neural versus Phrase-Based Machine Translation Quality: a Case Study

Within the field of Statistical Machine Translation (SMT), the neural ap...
research
09/03/2021

Language Modeling, Lexical Translation, Reordering: The Training Process of NMT through the Lens of Classical SMT

Differently from the traditional statistical MT that decomposes the tran...
research
01/04/2020

A Comprehensive Survey of Multilingual Neural Machine Translation

We present a survey on multilingual neural machine translation (MNMT), w...
research
05/01/2015

Hierarchy of Scales in Language Dynamics

Methods and insights from statistical physics are finding an increasing ...

Please sign up or login with your details

Forgot password? Click here to reset