Swin-Pose: Swin Transformer Based Human Pose Estimation

01/19/2022
by   Zinan Xiong, et al.
0

Convolutional neural networks (CNNs) have been widely utilized in many computer vision tasks. However, CNNs have a fixed reception field and lack the ability of long-range perception, which is crucial to human pose estimation. Due to its capability to capture long-range dependencies between pixels, transformer architecture has been adopted to computer vision applications recently and is proven to be a highly effective architecture. We are interested in exploring its capability in human pose estimation, and thus propose a novel model based on transformer architecture, enhanced with a feature pyramid fusion structure. More specifically, we use pre-trained Swin Transformer as our backbone and extract features from input images, we leverage a feature pyramid structure to extract feature maps from different stages. By fusing the features together, our model predicts the keypoint heatmap. The experiment results of our study have demonstrated that the proposed transformer-based model can achieve better performance compared to the state-of-the-art CNN-based models.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
03/24/2022

CrossFormer: Cross Spatio-Temporal Transformer for 3D Human Pose Estimation

3D human pose estimation can be handled by encoding the geometric depend...
research
09/21/2022

Sar Ship Detection based on Swin Transformer and Feature Enhancement Feature Pyramid Network

With the booming of Convolutional Neural Networks (CNNs), CNNs such as V...
research
04/13/2022

Recognition of Freely Selected Keypoints on Human Limbs

Nearly all Human Pose Estimation (HPE) datasets consist of a fixed set o...
research
05/08/2016

Chained Predictions Using Convolutional Neural Networks

In this paper, we present an adaptation of the sequence-to-sequence mode...
research
03/26/2021

Lifting Transformer for 3D Human Pose Estimation in Video

Despite great progress in video-based 3D human pose estimation, it is st...
research
07/29/2021

Efficient Human Pose Estimation by Maximizing Fusion and High-Level Spatial Attention

In this paper, we propose an efficient human pose estimation network – S...
research
06/07/2023

Efficient Vision Transformer for Human Pose Estimation via Patch Selection

While Convolutional Neural Networks (CNNs) have been widely successful i...

Please sign up or login with your details

Forgot password? Click here to reset