Spatial-Temporal Parallel Transformer for Arm-Hand Dynamic Estimation

03/30/2022
by   Shuying Liu, et al.
0

We propose an approach to estimate arm and hand dynamics from monocular video by utilizing the relationship between arm and hand. Although monocular full human motion capture technologies have made great progress in recent years, recovering accurate and plausible arm twists and hand gestures from in-the-wild videos still remains a challenge. To solve this problem, our solution is proposed based on the fact that arm poses and hand gestures are highly correlated in most real situations. To fully exploit arm-hand correlation as well as inter-frame information, we carefully design a Spatial-Temporal Parallel Arm-Hand Motion Transformer (PAHMT) to predict the arm and hand dynamics simultaneously. We also introduce new losses to encourage the estimations to be smooth and accurate. Besides, we collect a motion capture dataset including 200K frames of hand gestures and use this data to train our model. By integrating a 2D hand pose estimation model and a 3D human pose estimation model, the proposed method can produce plausible arm and hand dynamics from monocular video. Extensive evaluations demonstrate that the proposed method has advantages over previous state-of-the-art approaches and shows robustness under various challenging scenarios.

READ FULL TEXT

page 1

page 4

page 7

page 8

research
08/19/2020

FrankMocap: Fast Monocular 3D Hand and Body Motion Capture by Regression and Integration

Although the essential nuance of human motion is often conveyed as a com...
research
06/10/2019

Learning Individual Styles of Conversational Gesture

Human speech is often accompanied by hand and arm gestures. Given audio ...
research
07/23/2020

Body2Hands: Learning to Infer 3D Hands from Conversational Gesture Body Dynamics

We propose a novel learned deep prior of body motion for 3D hand shape s...
research
06/28/2021

Motion Projection Consistency Based 3D Human Pose Estimation with Virtual Bones from Monocular Videos

Real-time 3D human pose estimation is crucial for human-computer interac...
research
03/09/2023

Deformer: Dynamic Fusion Transformer for Robust Hand Pose Estimation

Accurately estimating 3D hand pose is crucial for understanding how huma...
research
10/11/2022

HiFECap: Monocular High-Fidelity and Expressive Capture of Human Performances

Monocular 3D human performance capture is indispensable for many applica...
research
01/08/2019

A Spatial-temporal 3D Human Pose Reconstruction Framework

3D human pose reconstruction from single-view camera is a difficult and ...

Please sign up or login with your details

Forgot password? Click here to reset