EFormer: Enhanced Transformer towards Semantic-Contour Features of Foreground for Portraits Matting

08/24/2023
by   Zitao Wang, et al.
0

The portrait matting task aims to extract an alpha matte with complete semantics and finely-detailed contours. In comparison to CNN-based approaches, transformers with self-attention allow a larger receptive field, enabling it to better capture long-range dependencies and low-frequency semantic information of a portrait. However, the recent research shows that self-attention mechanism struggle with modeling high-frequency information and capturing fine contour details, which can lead to bias while predicting the portrait's contours. To address the problem, we propose EFormer to enhance the model's attention towards semantic and contour features. Especially the latter, which is surrounded by a large amount of high-frequency details. We build a semantic and contour detector (SCD) to accurately capture the distribution of semantic and contour features. And we further design contour-edge extraction branch and semantic extraction branch for refining contour features and complete semantic information. Finally, we fuse the two kinds of features and leverage the segmentation head to generate the predicted portrait matte. Remarkably, EFormer is an end-to-end trimap-free method and boasts a simple structure. Experiments conducted on VideoMatte240K-JPEGSD and AIM datasets demonstrate that EFormer outperforms previous portrait matte methods.

READ FULL TEXT

page 2

page 12

page 13

research
04/18/2023

Frequency Enhanced Hybrid Attention Network for Sequential Recommendation

The self-attention mechanism, which equips with a strong capability of m...
research
12/02/2022

Dunhuang murals contour generation network based on convolution and self-attention fusion

Dunhuang murals are a collection of Chinese style and national style, fo...
research
03/21/2019

Context-Constrained Accurate Contour Extraction for Occlusion Edge Detection

Occlusion edge detection requires both accurate locations and context co...
research
08/31/2023

Laplacian-Former: Overcoming the Limitations of Vision Transformers in Local Texture Detection

Vision Transformer (ViT) models have demonstrated a breakthrough in a wi...
research
08/25/2023

Unlocking Fine-Grained Details with Wavelet-based High-Frequency Enhancement in Transformers

Medical image segmentation is a critical task that plays a vital role in...
research
11/24/2021

MorphMLP: A Self-Attention Free, MLP-Like Backbone for Image and Video

Self-attention has become an integral component of the recent network ar...
research
12/21/2017

Smart, Sparse Contours to Represent and Edit Images

We study the problem of reconstructing an image from information stored ...

Please sign up or login with your details

Forgot password? Click here to reset