Self-attention aggregation network for video face representation and recognition

10/11/2020
by   Ihor Protsenko, et al.
11

Models based on self-attention mechanisms have been successful in analyzing temporal data and have been widely used in the natural language domain. We propose a new model architecture for video face representation and recognition based on a self-attention mechanism. Our approach could be used for video with single and multiple identities. To the best of our knowledge, no one has explored the aggregation approaches that consider the video with multiple identities. The proposed approach utilizes existing models to get the face representation for each video frame, e.g., ArcFace and MobileFaceNet, and the aggregation module produces the aggregated face representation vector for video by taking into consideration the order of frames and their quality scores. We demonstrate empirical results on a public dataset for video face recognition called IJB-C to indicate that the self-attention aggregation network (SAAN) outperforms naive average pooling. Moreover, we introduce a new multi-identity video dataset based on the publicly available UMDFaces dataset and collected GIFs from Giphy. We show that SAAN is capable of producing a compact face representation for both single and multiple identities in a video. The dataset and source code will be publicly available.

READ FULL TEXT

page 4

page 5

research
05/06/2019

Fine-grained Attention-based Video Face Recognition

This paper aims to learn a compact representation of a video for video f...
research
11/20/2022

MINTIME: Multi-Identity Size-Invariant Video Deepfake Detection

In this paper, we introduce MINTIME, a video deepfake detection approach...
research
11/16/2020

Text Information Aggregation with Centrality Attention

A lot of natural language processing problems need to encode the text se...
research
04/26/2019

Recurrent Embedding Aggregation Network for Video Face Recognition

Recurrent networks have been successful in analyzing temporal data and h...
research
06/07/2021

Video Imprint

A new unified video analytics framework (ER3) is proposed for complex ev...
research
08/27/2023

FaceCoresetNet: Differentiable Coresets for Face Set Recognition

In set-based face recognition, we aim to compute the most discriminative...
research
09/22/2016

Pose-Selective Max Pooling for Measuring Similarity

In this paper, we deal with two challenges for measuring the similarity ...

Please sign up or login with your details

Forgot password? Click here to reset