Non-local convolutional neural networks (nlcnn) for speaker recognition

by   Haici Yang, et al.

Speaker recognition is the process of identifying a speaker based on the voice. The technology has attracted more attention with the recent increase in popularity of smart voice assistants, such as Amazon Alexa. In the past few years, various convolutional neural network (CNN) based speaker recognition algorithms have been proposed and achieved satisfactory performance. However, convolutional operations are building blocks that typically perform on a local neighborhood at a time and thus miss to capture global, long-range interactions at the feature level which are critical for understanding the pattern in a speaker's voice. In this work, we propose to apply Non-local Convolutional Neural Networks (NLCNN) to improve the capability of capturing long-range dependencies at the feature level, therefore improving speaker recognition performance. Specifically, we introduce non-local blocks where the output response of a position is computed as a weighted sum of the input features at all positions. Combining non-local blocks with pre-defined CNN networks, we investigate the effectiveness of NLCNN models. Without extensive tuning, the proposed NLCNN models outperform state-of-the-art speaker recognition algorithms on the public Voxceleb dataset. What's more, we investigate different types of non-local operations applied to the frequency-time domain, time domain, frequency domain and frame-level respectively. Among them, time domain is the most effective one for speaker recognition applications.


page 1

page 2

page 3

page 4


Non-local Neural Networks

Both convolutional and recurrent operations are building blocks that pro...

Region-based Non-local Operation for Video Classification

Convolutional Neural Networks (CNNs) model long-range dependencies by de...

Poly-NL: Linear Complexity Non-local Layers with Polynomials

Spatial self-attention layers, in the form of Non-Local blocks, introduc...

LSSANet: A Long Short Slice-Aware Network for Pulmonary Nodule Detection

Convolutional neural networks (CNNs) have been demonstrated to be highly...

Non-local Recurrent Neural Memory for Supervised Sequence Modeling

Typical methods for supervised sequence modeling are built upon the recu...

AttentionNAS: Spatiotemporal Attention Cell Search for Video Classification

Convolutional operations have two limitations: (1) do not explicitly mod...

Qiniu Submission to ActivityNet Challenge 2018

In this paper, we introduce our submissions for the tasks of trimmed act...

Please sign up or login with your details

Forgot password? Click here to reset