Unsupervised Cross-Modal Alignment for Multi-Person 3D Pose Estimation

08/04/2020
by   Jogendra Nath Kundu, et al.
1

We present a deployment friendly, fast bottom-up framework for multi-person 3D human pose estimation. We adopt a novel neural representation of multi-person 3D pose which unifies the position of person instances with their corresponding 3D pose representation. This is realized by learning a generative pose embedding which not only ensures plausible 3D pose predictions, but also eliminates the usual keypoint grouping operation as employed in prior bottom-up approaches. Further, we propose a practical deployment paradigm where paired 2D or 3D pose annotations are unavailable. In the absence of any paired supervision, we leverage a frozen network, as a teacher model, which is trained on an auxiliary task of multi-person 2D pose estimation. We cast the learning as a cross-modal alignment problem and propose training objectives to realize a shared latent space between two diverse modalities. We aim to enhance the model's ability to perform beyond the limiting teacher network by enriching the latent-to-3D pose mapping using artificially synthesized multi-person 3D scene samples. Our approach not only generalizes to in-the-wild images, but also yields a superior trade-off between speed and performance, compared to prior top-down approaches. Our approach also yields state-of-the-art multi-person 3D pose estimation performance among the bottom-up approaches under consistent supervision levels.

READ FULL TEXT

page 13

page 16

page 20

page 22

page 23

page 24

page 25

page 26

research
07/16/2020

Self-supervision on Unlabelled OR Data for Multi-person 2D/3D Human Pose Estimation

2D/3D human pose estimation is needed to develop novel intelligent tools...
research
04/01/2020

Knowledge as Priors: Cross-Modal Knowledge Generalization for Datasets without Superior Knowledge

Cross-modal knowledge distillation deals with transferring knowledge fro...
research
04/19/2021

LaLaLoc: Latent Layout Localisation in Dynamic, Unvisited Environments

We present LaLaLoc to localise in environments without the need for prio...
research
11/08/2018

Improving Multi-Person Pose Estimation using Label Correction

Significant attention is being paid to multi-person pose estimation meth...
research
04/24/2020

Neural Head Reenactment with Latent Pose Descriptors

We propose a neural head reenactment system, which is driven by a latent...
research
04/05/2022

Non-Local Latent Relation Distillation for Self-Adaptive 3D Human Pose Estimation

Available 3D human pose estimation approaches leverage different forms o...
research
04/04/2022

Aligning Silhouette Topology for Self-Adaptive 3D Human Pose Recovery

Articulation-centric 2D/3D pose supervision forms the core training obje...

Please sign up or login with your details

Forgot password? Click here to reset