Domain Adaptation of low-resource Target-Domain models using well-trained ASR Conformer Models

02/18/2022
by   Vrunda N. Sukhadia, et al.
0

In this paper, we investigate domain adaptation for low-resource Automatic Speech Recognition (ASR) of target-domain data, when a well-trained ASR model trained with a large dataset is available. We argue that in the encoder-decoder framework, the decoder of the well-trained ASR model is largely tuned towards the source-domain, hurting the performance of target-domain models in vanilla transfer-learning. On the other hand, the encoder layers of the well-trained ASR model mostly capture the acoustic characteristics. We, therefore, propose to use the embeddings tapped from these encoder layers as features for a downstream Conformer target-domain model and show that they provide significant improvements. We do ablation studies on which encoder layer is optimal to tap the embeddings, as well as the effect of freezing or updating the well-trained ASR model's encoder layers. We further show that applying Spectral Augmentation (SpecAug) on the proposed features (this is in addition to default SpecAug on input spectral features) provides a further improvement on the target-domain performance. For the LibriSpeech-100-clean data as target-domain and SPGI-5000 as a well-trained model, we get 30 Similarly, with WSJ data as target-domain and LibriSpeech-960 as a well-trained model, we get 50

READ FULL TEXT
research
09/12/2021

Unsupervised Domain Adaptation Schemes for Building ASR in Low-resource Languages

Building an automatic speech recognition (ASR) system from scratch requi...
research
01/24/2020

Data Techniques For Online End-to-end Speech Recognition

Practitioners often need to build ASR systems for new use cases in a sho...
research
08/25/2023

Decoupled Structure for Improved Adaptability of End-to-End Models

Although end-to-end (E2E) trainable automatic speech recognition (ASR) h...
research
10/08/2020

Gender domain adaptation for automatic speech recognition task

This paper is focused on the finetuning of acoustic models for speaker a...
research
05/07/2021

Self-Adaptive Transfer Learning for Multicenter Glaucoma Classification in Fundus Retina Images

The early diagnosis and screening of glaucoma are important for patients...
research
06/01/2023

Adapting an Unadaptable ASR System

As speech recognition model sizes and training data requirements grow, i...
research
10/08/2021

A Study of Low-Resource Speech Commands Recognition based on Adversarial Reprogramming

In this study, we propose a novel adversarial reprogramming (AR) approac...

Please sign up or login with your details

Forgot password? Click here to reset