Machine Translation Customization via Automatic Training Data Selection from the Web

02/20/2021
by   Thuy Vu, et al.
7

Machine translation (MT) systems, especially when designed for an industrial setting, are trained with general parallel data derived from the Web. Thus, their style is typically driven by word/structure distribution coming from the average of many domains. In contrast, MT customers want translations to be specialized to their domain, for which they are typically able to provide text samples. We describe an approach for customizing MT systems on specific domains by selecting data similar to the target customer data to train neural translation models. We build document classifiers using monolingual target data, e.g., provided by the customers to select parallel training data from Web crawled data. Finally, we train MT models on our automatically selected data, obtaining a system specialized to the target domain. We tested our approach on the benchmark from WMT-18 Translation Task for News domains enabling comparisons with state-of-the-art MT systems. The results show that our models outperform the top systems while using less data and smaller models.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
03/19/2019

CVIT-MT Systems for WAT-2018

This document describes the machine translation system used in the submi...
research
06/19/2019

The Effect of Translationese in Machine Translation Test Sets

The effect of translationese has been studied in the field of machine tr...
research
10/31/2019

Naver Labs Europe's Systems for the Document-Level Generation and Translation Task at WNGT 2019

Recently, neural models led to significant improvements in both machine ...
research
09/10/2021

Dynamic Terminology Integration for COVID-19 and other Emerging Domains

The majority of language domains require prudent use of terminology to e...
research
05/08/2023

Target-Side Augmentation for Document-Level Machine Translation

Document-level machine translation faces the challenge of data sparsity ...
research
08/11/2022

Domain-Specific Text Generation for Machine Translation

Preservation of domain knowledge from the source to target is crucial in...
research
12/14/2016

Unsupervised Clustering of Commercial Domains for Adaptive Machine Translation

In this paper, we report on domain clustering in the ambit of an adaptiv...

Please sign up or login with your details

Forgot password? Click here to reset