Layered Controllable Video Generation

11/24/2021
by   Jiahui Huang, et al.
0

We introduce layered controllable video generation, where we, without any supervision, decompose the initial frame of a video into foreground and background layers, with which the user can control the video generation process by simply manipulating the foreground mask. The key challenges are the unsupervised foreground-background separation, which is ambiguous, and ability to anticipate user manipulations with access to only raw video sequences. We address these challenges by proposing a two-stage learning procedure. In the first stage, with the rich set of losses and dynamic foreground size prior, we learn how to separate the frame into foreground and background layers and, conditioned on these layers, how to generate the next frame using VQ-VAE generator. In the second stage, we fine-tune this network to anticipate edits to the mask, by fitting (parameterized) control to the mask from future frame. We demonstrate the effectiveness of this learning and the more granular control mechanism, while illustrating state-of-the-art performance on two benchmark datasets. We provide a video abstract as well as some video results on https://gabriel-huang.github.io/layered_controllable_video_generation

READ FULL TEXT

page 3

page 7

page 8

page 13

page 14

page 15

research
03/26/2022

V3GAN: Decomposing Background, Foreground and Motion for Video Generation

Video generation is a challenging task that requires modeling plausible ...
research
04/21/2021

Shadow Generation for Composite Image in Real-world Scenes

Image composition targets at inserting a foreground object on a backgrou...
research
11/19/2021

Xp-GAN: Unsupervised Multi-object Controllable Video Generation

Video Generation is a relatively new and yet popular subject in machine ...
research
10/24/2019

Controllable Attention for Structured Layered Video Decomposition

The objective of this paper is to be able to separate a video into its n...
research
12/23/2021

Iteratively Selecting an Easy Reference Frame Makes Unsupervised Video Object Segmentation Easier

Unsupervised video object segmentation (UVOS) is a per-pixel binary labe...
research
12/20/2022

Video Segmentation Learning Using Cascade Residual Convolutional Neural Network

Video segmentation consists of a frame-by-frame selection process of mea...
research
02/09/2017

L1-regularized Reconstruction Error as Alpha Matte

Sampling-based alpha matting methods have traditionally followed the com...

Please sign up or login with your details

Forgot password? Click here to reset