Improving Ultrasound Tongue Image Reconstruction from Lip Images Using Self-supervised Learning and Attention Mechanism

06/20/2021
by   Haiyang Liu, et al.
0

Speech production is a dynamic procedure, which involved multi human organs including the tongue, jaw and lips. Modeling the dynamics of the vocal tract deformation is a fundamental problem to understand the speech, which is the most common way for human daily communication. Researchers employ several sensory streams to describe the process simultaneously, which are incontrovertibly statistically related to other streams. In this paper, we address the following question: given an observable image sequences of lips, can we picture the corresponding tongue motion. We formulated this problem as the self-supervised learning problem, and employ the two-stream convolutional network and long-short memory network for the learning task, with the attention mechanism. We evaluate the performance of the proposed method by leveraging the unlabeled lip videos to predict an upcoming ultrasound tongue image sequence. The results show that our model is able to generate images that close to the real ultrasound tongue images, and results in the matching between two imaging modalities.

READ FULL TEXT

page 3

page 4

research
04/27/2020

Self-Supervised Attention Learning for Depth and Ego-motion Estimation

We address the problem of depth and ego-motion estimation from image seq...
research
02/19/2019

Predicting tongue motion in unlabeled ultrasound videos using convolutional LSTM neural network

A challenge in speech production research is to predict future tongue mo...
research
01/27/2021

Convolutional Neural Network-Based Age Estimation Using B-Mode Ultrasound Tongue Image

Ultrasound tongue imaging is widely used for speech production research,...
research
05/31/2021

Automatic audiovisual synchronisation for ultrasound tongue imaging

Ultrasound tongue imaging is used to visualise the intra-oral articulato...
research
06/15/2022

Self-Supervised Implicit Attention: Guided Attention by The Model Itself

We propose Self-Supervised Implicit Attention (SSIA), a new approach tha...
research
04/15/2018

Attention-Gated Networks for Improving Ultrasound Scan Plane Detection

In this work, we apply an attention-gated network to real-time automated...
research
08/26/2023

A small vocabulary database of ultrasound image sequences of vocal tract dynamics

This paper presents a new database consisting of concurrent articulatory...

Please sign up or login with your details

Forgot password? Click here to reset