Word Embeddings for Automatic Equalization in Audio Mixing

02/17/2022
by   Satvik Venkatesh, et al.
7

In recent years, machine learning has been widely adopted to automate the audio mixing process. Automatic mixing systems have been applied to various audio effects such as gain-adjustment, stereo panning, equalization, and reverberation. These systems can be controlled through visual interfaces, providing audio examples, using knobs, and semantic descriptors. Using semantic descriptors or textual information to control these systems is an effective way for artists to communicate their creative goals. Furthermore, sometimes artists use non-technical words that may not be understood by the mixing system, or even a mixing engineer. In this paper, we explore the novel idea of using word embeddings to represent semantic descriptors. Word embeddings are generally obtained by training neural networks on large corpora of written text. These embeddings serve as the input layer of the neural network to create a translation from words to EQ settings. Using this technique, the machine learning model can also generate EQ settings for semantic descriptors that it has not seen before. We perform experiments to demonstrate the feasibility of this idea. In addition, we compare the EQ settings of humans with the predictions of the neural network to evaluate the quality of predictions. The results showed that the embedding layer enables the neural network to understand semantic descriptors. We observed that the models with embedding layers perform better those without embedding layers, but not as good as human labels.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
03/23/2018

Speech2Vec: A Sequence-to-Sequence Framework for Learning Word Embeddings from Speech

In this paper, we propose a novel deep neural network architecture, Spee...
research
05/16/2020

RPD: A Distance Function Between Word Embeddings

It is well-understood that different algorithms, training processes, and...
research
04/18/2020

Effect of Text Color on Word Embeddings

In natural scenes and documents, we can find the correlation between a t...
research
04/04/2019

ReWE: Regressing Word Embeddings for Regularization of Neural Machine Translation Systems

Regularization of neural machine translation is still a significant prob...
research
10/19/2021

Inter-Sense: An Investigation of Sensory Blending in Fiction

This study reports on the semantic organization of English sensory descr...
research
10/20/2020

Automatic multitrack mixing with a differentiable mixing console of neural audio effects

Applications of deep learning to automatic multitrack mixing are largely...
research
06/08/2022

Words are all you need? Capturing human sensory similarity with textual descriptors

Recent advances in multimodal training use textual descriptions to signi...

Please sign up or login with your details

Forgot password? Click here to reset