MP3net: coherent, minute-long music generation from raw audio with a simple convolutional GAN

01/12/2021
by   Korneel van den Broek, et al.
9

We present a deep convolutional GAN which leverages techniques from MP3/Vorbis audio compression to produce long, high-quality audio samples with long-range coherence. The model uses a Modified Discrete Cosine Transform (MDCT) data representation, which includes all phase information. Phase generation is hence integral part of the model. We leverage the auditory masking and psychoacoustic perception limit of the human ear to widen the true distribution and stabilize the training process. The model architecture is a deep 2D convolutional network, where each subsequent generator model block increases the resolution along the time axis and adds a higher octave along the frequency axis. The deeper layers are connected with all parts of the output and have the context of the full track. This enables generation of samples which exhibit long-range coherence. We use MP3net to create 95s stereo tracks with a 22kHz sample rate after training for 250h on a single Cloud TPUv2. An additional benefit of the CNN-based model architecture is that generation of new songs is almost instantaneous.

READ FULL TEXT

page 2

page 9

research
06/26/2018

Conditioning Deep Generative Raw Audio Models for Structured Automatic Music

Existing automatic music generation approaches that feature deep learnin...
research
06/26/2018

The challenge of realistic music generation: modelling raw audio at scale

Realistic music generation is a challenging task. When building generati...
research
08/02/2017

Audio Super Resolution using Neural Networks

We introduce a new audio processing technique that increases the samplin...
research
06/04/2019

MelNet: A Generative Model for Audio in the Frequency Domain

Capturing high-level structure in audio waveforms is challenging because...
research
06/04/2021

Fre-GAN: Adversarial Frequency-consistent Audio Synthesis

Although recent works on neural vocoder have improved the quality of syn...
research
06/03/2016

Incorporating long-range consistency in CNN-based texture generation

Gatys et al. (2015) showed that pair-wise products of features in a conv...
research
06/25/2021

Basis-MelGAN: Efficient Neural Vocoder Based on Audio Decomposition

Recent studies have shown that neural vocoders based on generative adver...

Please sign up or login with your details

Forgot password? Click here to reset