Selective Spatio-Temporal Aggregation Based Pose Refinement System: Towards Understanding Human Activities in Real-World Videos

11/10/2020
by   Di Yang, et al.
0

Taking advantage of human pose data for understanding human activities has attracted much attention these days. However, state-of-the-art pose estimators struggle in obtaining high-quality 2D or 3D pose data due to occlusion, truncation and low-resolution in real-world un-annotated videos. Hence, in this work, we propose 1) a Selective Spatio-Temporal Aggregation mechanism, named SST-A, that refines and smooths the keypoint locations extracted by multiple expert pose estimators, 2) an effective weakly-supervised self-training framework which leverages the aggregated poses as pseudo ground-truth instead of handcrafted annotations for real-world pose estimation. Extensive experiments are conducted for evaluating not only the upstream pose refinement but also the downstream action recognition performance on four datasets, Toyota Smarthome, NTU-RGB+D, Charades, and Kinetics-50. We demonstrate that the skeleton data refined by our Pose-Refinement system (SSTA-PRS) is effective at boosting various existing action recognition models, which achieves competitive or state-of-the-art performance.

READ FULL TEXT

page 1

page 3

page 6

research
02/24/2021

PFRL: Pose-Free Reinforcement Learning for 6D Pose Estimation

6D pose estimation from a single RGB image is a challenging and vital ta...
research
07/24/2021

TinyAction Challenge: Recognizing Real-world Low-resolution Activities in Videos

This paper summarizes the TinyAction challenge which was organized in Ac...
research
07/19/2021

UNIK: A Unified Framework for Real-world Skeleton-based Action Recognition

Action recognition based on skeleton data has recently witnessed increas...
research
02/01/2022

ADG-Pose: Automated Dataset Generation for Real-World Human Pose Estimation

Recent advancements in computer vision have seen a rise in the prominenc...
research
09/20/2022

Hierarchical Temporal Transformer for 3D Hand Pose Estimation and Action Recognition from Egocentric RGB Videos

Understanding dynamic hand motions and actions from egocentric RGB video...
research
12/27/2021

SmoothNet: A Plug-and-Play Network for Refining Human Poses in Videos

When analyzing human motion videos, the output jitters from existing pos...

Please sign up or login with your details

Forgot password? Click here to reset