Spatiotemporal Modeling for Crowd Counting in Videos

07/25/2017
by   Feng Xiong, et al.
0

Region of Interest (ROI) crowd counting can be formulated as a regression problem of learning a mapping from an image or a video frame to a crowd density map. Recently, convolutional neural network (CNN) models have achieved promising results for crowd counting. However, even when dealing with video data, CNN-based methods still consider each video frame independently, ignoring the strong temporal correlation between neighboring frames. To exploit the otherwise very useful temporal information in video sequences, we propose a variant of a recent deep learning model called convolutional LSTM (ConvLSTM) for crowd counting. Unlike the previous CNN-based methods, our method fully captures both spatial and temporal dependencies. Furthermore, we extend the ConvLSTM model to a bidirectional ConvLSTM model which can access long-range information in both directions. Extensive experiments using four publicly available datasets demonstrate the reliability of our approach and the effectiveness of incorporating temporal information to boost the accuracy of crowd counting. In addition, we also conduct some transfer learning experiments to show that once our model is trained on one dataset, its learning experience can be transferred easily to a new dataset which consists of only very few video frames for model adaptation.

READ FULL TEXT

page 5

page 6

page 7

page 8

research
07/18/2019

Locality-constrained Spatial Transformer Network for Video Crowd Counting

Compared with single image based crowd counting, video provides the spat...
research
08/12/2019

Enhanced 3D convolutional networks for crowd counting

Recently, convolutional neural networks (CNNs) are the leading defacto m...
research
07/04/2019

Video Crowd Counting via Dynamic Temporal Modeling

Crowd counting aims to count the number of instantaneous people in a cro...
research
07/13/2021

Developmental Stage Classification of Embryos Using Two-Stream Neural Network with Linear-Chain Conditional Random Field

The developmental process of embryos follows a monotonic order. An embry...
research
11/28/2018

Non-Volume Preserving-based Feature Fusion Approach to Group-Level Expression Recognition on Crowd Videos

Group-level emotion recognition (ER) is a growing research area as the d...
research
05/15/2018

A Deeply-Recursive Convolutional Network for Crowd Counting

The estimation of crowd count in images has a wide range of applications...
research
02/15/2019

TMAV: Temporal Motionless Analysis of Video using CNN in MPSoC

Analyzing video for traffic categorization is an important pillar of Int...

Please sign up or login with your details

Forgot password? Click here to reset