Modulated Variational auto-Encoders for many-to-many musical timbre transfer

09/29/2018
by   Adrien Bitton, et al.
0

Generative models have been successfully applied to image style transfer and domain translation. However, there is still a wide gap in the quality of results when learning such tasks on musical audio. Furthermore, most translation models only enable one-to-one or one-to-many transfer by relying on separate encoders or decoders and complex, computationally-heavy models. In this paper, we introduce the Modulated Variational auto-Encoders (MoVE) to perform musical timbre transfer. We define timbre transfer as applying parts of the auditory properties of a musical instrument onto another. First, we show that we can achieve this task by conditioning existing domain translation techniques with Feature-wise Linear Modulation (FiLM). Then, we alleviate the need for additional adversarial networks by replacing the usual translation criterion by a Maximum Mean Discrepancy (MMD) objective. This allows a faster and more stable training along with a controllable latent space encoder. By further conditioning our system on several different instruments, we can generalize to many-to-many transfer within a single variational architecture able to perform multi-domain transfers. Our models map inputs to 3-dimensional representations, successfully translating timbre from one instrument to another and supporting sound synthesis from a reduced set of control parameters. We evaluate our method in reconstruction and generation tasks while analyzing the auditory descriptor distributions across transferred domains. We show that this architecture allows for generative controls in multi-domain transfer, yet remaining light, fast to train and effective on small datasets.

READ FULL TEXT
research
05/30/2019

Musical Composition Style Transfer via Disentangled Timbre Representations

Music creation involves not only composing the different parts (e.g., me...
research
04/12/2019

Assisted Sound Sample Generation with Musical Conditioning in Adversarial Auto-Encoders

Generative models have thrived in computer vision, enabling unprecedente...
research
09/05/2021

Timbre Transfer with Variational Auto Encoding and Cycle-Consistent Adversarial Networks

This research project investigates the application of deep learning to t...
research
06/19/2019

Learning Disentangled Representations of Timbre and Pitch for Musical Instrument Sounds Using Gaussian Mixture Variational Autoencoders

In this paper, we learn disentangled representations of timbre and pitch...
research
02/27/2023

Continuous descriptor-based control for deep audio synthesis

Despite significant advances in deep models for music generation, the us...
research
05/22/2018

Generative timbre spaces: regularizing variational auto-encoders with perceptual metrics

Timbre spaces have been used in music perception to study the perceptual...
research
02/10/2020

Cross-modal variational inference for bijective signal-symbol translation

Extraction of symbolic information from signals is an active field of re...

Please sign up or login with your details

Forgot password? Click here to reset