β-Multivariational Autoencoder for Entangled Representation Learning in Video Frames

11/22/2022
by   Fatemeh Nouri, et al.
0

It is crucial to choose actions from an appropriate distribution while learning a sequential decision-making process in which a set of actions is expected given the states and previous reward. Yet, if there are more than two latent variables and every two variables have a covariance value, learning a known prior from data becomes challenging. Because when the data are big and diverse, many posterior estimate methods experience posterior collapse. In this paper, we propose the β-Multivariational Autoencoder (βMVAE) to learn a Multivariate Gaussian prior from video frames for use as part of a single object-tracking in form of a decision-making process. We present a novel formulation for object motion in videos with a set of dependent parameters to address a single object-tracking task. The true values of the motion parameters are obtained through data analysis on the training set. The parameters population is then assumed to have a Multivariate Gaussian distribution. The βMVAE is developed to learn this entangled prior p = N(μ, Σ) directly from frame patches where the output is the object masks of the frame patches. We devise a bottleneck to estimate the posterior's parameters, i.e. μ', Σ'. Via a new reparameterization trick, we learn the likelihood p(x̂|z) as the object mask of the input. Furthermore, we alter the neural network of βMVAE with the U-Net architecture and name the new network βMultivariational U-Net (βMVUnet). Our networks are trained from scratch via over 85k video frames for 24 (βMVUnet) and 78 (βMVAE) million steps. We show that βMVUnet enhances both posterior estimation and segmentation functioning over the test set. Our code and the trained networks are publicly released.

READ FULL TEXT

page 11

page 12

page 13

page 15

page 16

page 23

research
07/05/2023

ZJU ReLER Submission for EPIC-KITCHEN Challenge 2023: TREK-150 Single Object Tracking

The Associating Objects with Transformers (AOT) framework has exhibited ...
research
01/01/2022

PatchTrack: Multiple Object Tracking Using Frame Patches

Object motion and object appearance are commonly used information in mul...
research
03/28/2017

Lucid Data Dreaming for Multiple Object Tracking

Convolutional networks reach top quality in pixel-level object tracking ...
research
01/06/2018

Learning Hierarchical Features for Visual Object Tracking with Recursive Neural Networks

Recently, deep learning has achieved very promising results in visual ob...
research
07/17/2017

Tracking as Online Decision-Making: Learning a Policy from Streaming Videos with Reinforcement Learning

We formulate tracking as an online decision-making process, where a trac...
research
03/15/2019

Inserting Videos into Videos

In this paper, we introduce a new problem of manipulating a given video ...

Please sign up or login with your details

Forgot password? Click here to reset