Deep Learning for Robust Motion Segmentation with Non-Static Cameras

02/22/2021
by   Markus Bosch, et al.
0

This work proposes a new end-to-end DCNN based approach for motion segmentation, especially for video sequences captured with such non-static cameras, called MOSNET. While other approaches focus on spatial or temporal context only, the proposed approach uses 3D convolutions as a key technology to factor in, spatio-temporal features in cohesive video frames. This is done by capturing temporal information in features with a low and also with a high level of abstraction. The lean network architecture with about 21k trainable parameters is mainly based on a pre-trained VGG-16 network. The MOSNET uses a new feature map fusion technique, which enables the network to focus on the appropriate level of abstraction, resolution, and the appropriate size of the receptive field regarding the input. Furthermore, the end-to-end deep learning based approach can be extended by feature based image alignment as a pre-processing step, which brings a gain in performance for some scenes. Evaluating the end-to-end deep learning based MOSNET network in a scene independent manner leads to an overall F-measure of 0.803 on the CDNet2014 dataset. A small temporal window of five input frames, without the need of any initialization is used to obtain this result. Therefore the network is able to perform well on scenes captured with non-static cameras where the image content changes significantly during the scene. In order to get robust results in scenes captured with a moving camera, feature based image alignment can implemented as pre-processing step. The MOSNET combined with pre-processing leads to an F-measure of 0.685 when cross-evaluating with a relabeled LASIESTA dataset, which underpins the capability generalise of the MOSNET.

READ FULL TEXT

page 24

page 30

page 40

research
11/27/2022

GRelPose: Generalizable End-to-End Relative Camera Pose Regression

This paper proposes a generalizable, end-to-end deep learning-based meth...
research
10/22/2020

Learning to Sort Image Sequences via Accumulated Temporal Differences

Consider a set of n images of a scene with dynamic objects captured with...
research
09/25/2016

Deep learning based fence segmentation and removal from an image using a video sequence

Conventional approaches to image de-fencing use multiple adjacent frames...
research
04/06/2023

Deep learning-based image exposure enhancement as a pre-processing for an accurate 3D colon surface reconstruction

This contribution shows how an appropriate image pre-processing can impr...
research
11/17/2022

TempNet: Temporal Attention Towards the Detection of Animal Behaviour in Videos

Recent advancements in cabled ocean observatories have increased the qua...
research
12/10/2016

Towards an Automated Image De-fencing Algorithm Using Sparsity

Conventional approaches to image de-fencing suffer from non-robust fence...
research
03/28/2018

Robust Video Content Alignment and Compensation for Rain Removal in a CNN Framework

Rain removal is important for improving the robustness of outdoor vision...

Please sign up or login with your details

Forgot password? Click here to reset