LibriS2S: A German-English Speech-to-Speech Translation Corpus

04/22/2022
by   Pedro Jeuris, et al.
0

Recently, we have seen an increasing interest in the area of speech-to-text translation. This has led to astonishing improvements in this area. In contrast, the activities in the area of speech-to-speech translation is still limited, although it is essential to overcome the language barrier. We believe that one of the limiting factors is the availability of appropriate training data. We address this issue by creating LibriS2S, to our knowledge the first publicly available speech-to-speech training corpus between German and English. For this corpus, we used independently created audio for German and English leading to an unbiased pronunciation of the text in both languages. This allows the creation of a new text-to-speech and speech-to-speech translation model that directly learns to generate the speech signal based on the pronunciation of the source language. Using this created corpus, we propose Text-to-Speech models based on the example of the recently proposed FastSpeech 2 model that integrates source language information. We do this by adapting the model to take information such as the pitch, energy or transcript from the source speech as additional input.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
10/17/2019

LibriVoxDeEn: A Corpus for German-to-English Speech Translation and Speech Recognition

We present a corpus of sentence-aligned triples of German audio, German ...
research
09/13/2017

Linguistic Features of Genre and Method Variation in Translation: A Computational Perspective

In this paper we describe the use of text classification methods to inve...
research
06/11/2021

Sprachsynthese – State-of-the-Art in englischer und deutscher Sprache

Reading text aloud is an important feature for modern computer applicati...
research
10/15/2021

Scribosermo: Fast Speech-to-Text models for German and other Languages

Recent Speech-to-Text models often require a large amount of hardware re...
research
10/09/2019

Spoken Language Identification using ConvNets

Language Identification (LI) is an important first step in several speec...
research
06/28/2022

On the Impact of Noises in Crowd-Sourced Data for Speech Translation

Training speech translation (ST) models requires large and high-quality ...
research
02/02/2021

CTC-based Compression for Direct Speech Translation

Previous studies demonstrated that a dynamic phone-informed compression ...

Please sign up or login with your details

Forgot password? Click here to reset