Category-Aware Transformer Network for Better Human-Object Interaction Detection

04/11/2022
by   Leizhen Dong, et al.
0

Human-Object Interactions (HOI) detection, which aims to localize a human and a relevant object while recognizing their interaction, is crucial for understanding a still image. Recently, transformer-based models have significantly advanced the progress of HOI detection. However, the capability of these models has not been fully explored since the Object Query of the model is always simply initialized as just zeros, which would affect the performance. In this paper, we try to study the issue of promoting transformer-based HOI detectors by initializing the Object Query with category-aware semantic information. To this end, we innovatively propose the Category-Aware Transformer Network (CATN). Specifically, the Object Query would be initialized via category priors represented by an external object detection model to yield better performance. Moreover, such category priors can be further used for enhancing the representation ability of features via the attention mechanism. We have firstly verified our idea via the Oracle experiment by initializing the Object Query with the groundtruth category information. And then extensive experiments have been conducted to show that a HOI detection model equipped with our idea outperforms the baseline by a large margin to achieve a new state-of-the-art result.

READ FULL TEXT
research
03/24/2023

Category Query Learning for Human-Object Interaction Classification

Unlike most previous HOI methods that focus on learning better human-obj...
research
06/13/2022

Exploring Structure-aware Transformer over Interaction Proposals for Human-Object Interaction Detection

Recent high-performing Human-Object Interaction (HOI) detection techniqu...
research
06/06/2021

Oriented Object Detection with Transformer

Object detection with Transformers (DETR) has achieved a competitive per...
research
01/31/2023

Priors are Powerful: Improving a Transformer for Multi-camera 3D Detection with 2D Priors

Transfomer-based approaches advance the recent development of multi-came...
research
07/19/2023

Mining Conditional Part Semantics with Occluded Extrapolation for Human-Object Interaction Detection

Human-Object Interaction Detection is a crucial aspect of human-centric ...
research
05/27/2017

CASENet: Deep Category-Aware Semantic Edge Detection

Boundary and edge cues are highly beneficial in improving a wide variety...
research
09/02/2023

S^3-MonoDETR: Supervised Shape Scale-perceptive Deformable Transformer for Monocular 3D Object Detection

Recently, transformer-based methods have shown exceptional performance i...

Please sign up or login with your details

Forgot password? Click here to reset