A novel multimodal dynamic fusion network for disfluency detection in spoken utterances

11/27/2022
by   Sreyan Ghosh, et al.
0

Disfluency, though originating from human spoken utterances, is primarily studied as a uni-modal text-based Natural Language Processing (NLP) task. Based on early-fusion and self-attention-based multimodal interaction between text and acoustic modalities, in this paper, we propose a novel multimodal architecture for disfluency detection from individual utterances. Our architecture leverages a multimodal dynamic fusion network that adds minimal parameters over an existing text encoder commonly used in prior art to leverage the prosodic and acoustic cues hidden in speech. Through experiments, we show that our proposed model achieves state-of-the-art results on the widely used English Switchboard for disfluency detection and outperforms prior unimodal and multimodal systems in literature by a significant margin. In addition, we make a thorough qualitative analysis and show that, unlike text-only systems, which suffer from spurious correlations in the data, our system overcomes this problem through additional cues from speech signals. We make all our codes publicly available on GitHub.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
03/31/2022

MMER: Multimodal Multi-task learning for Emotion Recognition in Spoken Utterances

Emotion Recognition (ER) aims to classify human utterances into differen...
research
10/14/2021

Speech Toxicity Analysis: A New Spoken Language Processing Task

Toxic speech, also known as hate speech, is regarded as one of the cruci...
research
03/30/2022

Span Classification with Structured Information for Disfluency Detection in Spoken Utterances

Existing approaches in disfluency detection focus on solving a token-lev...
research
06/05/2019

Towards Multimodal Sarcasm Detection (An _Obviously_ Perfect Paper)

Sarcasm is often expressed through several verbal and non-verbal cues, e...
research
04/02/2019

The Verbal and Non Verbal Signals of Depression -- Combining Acoustics, Text and Visuals for Estimating Depression Level

Depression is a serious medical condition that is suffered by a large nu...
research
06/05/2023

BeAts: Bengali Speech Acts Recognition using Multimodal Attention Fusion

Spoken languages often utilise intonation, rhythm, intensity, and struct...
research
04/08/2019

Giving Attention to the Unexpected: Using Prosody Innovations in Disfluency Detection

Disfluencies in spontaneous speech are known to be associated with proso...

Please sign up or login with your details

Forgot password? Click here to reset