MA-ViT: Modality-Agnostic Vision Transformers for Face Anti-Spoofing

04/15/2023
by   Ajian Liu, et al.
0

The existing multi-modal face anti-spoofing (FAS) frameworks are designed based on two strategies: halfway and late fusion. However, the former requires test modalities consistent with the training input, which seriously limits its deployment scenarios. And the latter is built on multiple branches to process different modalities independently, which limits their use in applications with low memory or fast execution requirements. In this work, we present a single branch based Transformer framework, namely Modality-Agnostic Vision Transformer (MA-ViT), which aims to improve the performance of arbitrary modal attacks with the help of multi-modal data. Specifically, MA-ViT adopts the early fusion to aggregate all the available training modalities data and enables flexible testing of any given modal samples. Further, we develop the Modality-Agnostic Transformer Block (MATB) in MA-ViT, which consists of two stacked attentions named Modal-Disentangle Attention (MDA) and Cross-Modal Attention (CMA), to eliminate modality-related information for each modal sequences and supplement modality-agnostic liveness features from another modal sequences, respectively. Experiments demonstrate that the single model trained based on MA-ViT can not only flexibly evaluate different modal samples, but also outperforms existing single-modal frameworks by a large margin, and approaches the multi-modal frameworks introduced with smaller FLOPs and model parameters.

READ FULL TEXT
research
09/30/2022

Husformer: A Multi-Modal Transformer for Multi-Modal Human State Recognition

Human state recognition is a critical topic with pervasive and important...
research
04/24/2020

PipeNet: Selective Modal Pipeline of Fusion Network for Multi-Modal Face Anti-Spoofing

Face anti-spoofing has become an increasingly important and critical sec...
research
02/16/2022

Flexible-Modal Face Anti-Spoofing: A Benchmark

Face anti-spoofing (FAS) plays a vital role in securing face recognition...
research
02/01/2023

Multispectral Pedestrian Detection via Reference Box Constrained Cross Attention and Modality Balanced Optimization

Multispectral pedestrian detection is an important task for many around-...
research
05/03/2023

SeqAug: Sequential Feature Resampling as a modality agnostic augmentation method

Data augmentation is a prevalent technique for improving performance in ...
research
10/01/2022

Cascaded Multi-Modal Mixing Transformers for Alzheimer's Disease Classification with Incomplete Data

Accurate medical classification requires a large number of multi-modal d...
research
12/29/2022

MEAformer: Multi-modal Entity Alignment Transformer for Meta Modality Hybrid

As an important variant of entity alignment (EA), multi-modal entity ali...

Please sign up or login with your details

Forgot password? Click here to reset