Audeo: Audio Generation for a Silent Performance Video

06/23/2020
by   Kun Su, et al.
0

We present a novel system that gets as an input video frames of a musician playing the piano and generates the music for that video. Generation of music from visual cues is a challenging problem and it is not clear whether it is an attainable goal at all. Our main aim in this work is to explore the plausibility of such a transformation and to identify cues and components able to carry the association of sounds with visual events. To achieve the transformation we built a full pipeline named `Audeo' containing three components. We first translate the video frames of the keyboard and the musician hand movements into raw mechanical musical symbolic representation Piano-Roll (Roll) for each video frame which represents the keys pressed at each time step. We then adapt the Roll to be amenable for audio synthesis by including temporal correlations. This step turns out to be critical for meaningful audio generation. As a last step, we implement Midi synthesizers to generate realistic music. Audeo converts video to audio smoothly and clearly with only a few setup constraints. We evaluate Audeo on `in the wild' piano performance videos and obtain that their generated music is of reasonable audio quality and can be successfully recognized with high precision by popular music identification software.

READ FULL TEXT

page 4

page 5

page 8

research
05/11/2023

V2Meow: Meowing to the Visual Beat via Music Generation

Generating high quality music that complements the visual content of a v...
research
12/07/2020

Multi-Instrumentalist Net: Unsupervised Generation of Music from Body Movements

We propose a novel system that takes as an input body movements of a mus...
research
07/21/2020

Foley Music: Learning to Generate Music from Videos

In this paper, we introduce Foley Music, a system that can synthesize pl...
research
04/01/2022

Quantized GAN for Complex Music Generation from Dance Videos

We present Dance2Music-GAN (D2M-GAN), a novel adversarial multi-modal fr...
research
03/25/2019

Learning Embodied Semantics via Music and Dance Semiotic Correlations

Music semantics is embodied, in the sense that meaning is biologically m...
research
12/19/2017

Audio to Body Dynamics

We present a method that gets as input an audio of violin or piano playi...
research
03/20/2017

Dance Dance Convolution

Dance Dance Revolution (DDR) is a popular rhythm-based video game. Playe...

Please sign up or login with your details

Forgot password? Click here to reset