SpeechYOLO: Detection and Localization of Speech Objects

04/14/2019
by   Yael Segal, et al.
0

In this paper, we propose to apply object detection methods from the vision domain on the speech recognition domain, by treating audio fragments as objects. More specifically, we present SpeechYOLO, which is inspired by the YOLO algorithm for object detection in images. The goal of SpeechYOLO is to localize boundaries of utterances within the input signal, and to correctly classify them. Our system is composed of a convolutional neural network, with a simple least-mean-squares loss function. We evaluated the system on several keyword spotting tasks, that include corpora of read speech and spontaneous speech. Our system compares favorably with other algorithms trained for both localization and classification.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
10/24/2022

I see what you hear: a vision-inspired method to localize words

This paper explores the possibility of using visual object detection tec...
research
10/18/2017

Honk: A PyTorch Reimplementation of Convolutional Neural Networks for Keyword Spotting

We describe Honk, an open-source PyTorch reimplementation of convolution...
research
06/13/2023

A Novel Scheme to classify Read and Spontaneous Speech

The COVID-19 pandemic has led to an increased use of remote telephonic i...
research
10/30/2022

Foreign Object Debris Detection for Airport Pavement Images based on Self-supervised Localization and Vision Transformer

Supervised object detection methods provide subpar performance when appl...
research
03/10/2018

Speech Recognition: Keyword Spotting Through Image Recognition

The problem of identifying voice commands has always been a challenge du...
research
08/31/2023

Improving vision-inspired keyword spotting using dynamic module skipping in streaming conformer encoder

Using a vision-inspired keyword spotting framework, we propose an archit...
research
08/23/2019

VOP Detection for Read and Conversation Speech using CWT Coefficients and Phone Boundaries

In this paper, we propose a novel approach for accurate detection of the...

Please sign up or login with your details

Forgot password? Click here to reset