Mode Variational LSTM Robust to Unseen Modes of Variation: Application to Facial Expression Recognition

11/16/2018
by   Wissam J. Baddar, et al.
0

Spatio-temporal feature encoding is essential for encoding the dynamics in video sequences. Recurrent neural networks, particularly long short-term memory (LSTM) units, have been popular as an efficient tool for encoding spatio-temporal features in sequences. In this work, we investigate the effect of mode variations on the encoded spatio-temporal features using LSTMs. We show that the LSTM retains information related to the mode variation in the sequence, which is irrelevant to the task at hand (e.g. classification facial expressions). Actually, the LSTM forget mechanism is not robust enough to mode variations and preserves information that could negatively affect the encoded spatio-temporal features. We propose the mode variational LSTM to encode spatio-temporal features robust to unseen modes of variation. The mode variational LSTM modifies the original LSTM structure by adding an additional cell state that focuses on encoding the mode variation in the input sequence. To efficiently regulate what features should be stored in the additional cell state, additional gating functionality is also introduced. The effectiveness of the proposed mode variational LSTM is verified using the facial expression recognition task. Comparative experiments on publicly available datasets verified that the proposed mode variational LSTM outperforms existing methods. Moreover, a new dynamic facial expression dataset with different modes of variation, including various modes like pose and illumination variations, was collected to comprehensively evaluate the proposed mode variational LSTM. Experimental results verified that the proposed mode variational LSTM encodes spatio-temporal features robust to unseen modes of variation.

READ FULL TEXT

page 3

page 5

research
11/29/2017

Learning Spatio-temporal Features with Partial Expression Sequences for on-the-Fly Prediction

Spatio-temporal feature encoding is essential for encoding facial expres...
research
05/10/2022

Spatio-Temporal Transformer for Dynamic Facial Expression Recognition in the Wild

Previous methods for dynamic facial expression in the wild are mainly ba...
research
10/26/2020

Video-based Facial Expression Recognition using Graph Convolutional Networks

Facial expression recognition (FER), aiming to classify the expression p...
research
05/29/2018

Microscopy Cell Segmentation via Convolutional LSTM Networks

Live cell microscopy sequences exhibit complex spatial structures and co...
research
04/11/2018

Deep Differential Recurrent Neural Networks

Due to the special gating schemes of Long Short-Term Memory (LSTM), LSTM...
research
09/19/2017

Reducing Complexity of HEVC: A Deep Learning Approach

High Efficiency Video Coding (HEVC) significantly reduces bit-rates over...

Please sign up or login with your details

Forgot password? Click here to reset