PPR-FCN: Weakly Supervised Visual Relation Detection via Parallel Pairwise R-FCN

08/07/2017
by   Hanwang Zhang, et al.
0

We aim to tackle a novel vision task called Weakly Supervised Visual Relation Detection (WSVRD) to detect "subject-predicate-object" relations in an image with object relation groundtruths available only at the image level. This is motivated by the fact that it is extremely expensive to label the combinatorial relations between objects at the instance level. Compared to the extensively studied problem, Weakly Supervised Object Detection (WSOD), WSVRD is more challenging as it needs to examine a large set of regions pairs, which is computationally prohibitive and more likely stuck in a local optimal solution such as those involving wrong spatial context. To this end, we present a Parallel, Pairwise Region-based, Fully Convolutional Network (PPR-FCN) for WSVRD. It uses a parallel FCN architecture that simultaneously performs pair selection and classification of single regions and region pairs for object and relation detection, while sharing almost all computation shared over the entire image. In particular, we propose a novel position-role-sensitive score map with pairwise RoI pooling to efficiently capture the crucial context associated with a pair of objects. We demonstrate the superiority of PPR-FCN over all baselines in solving the WSVRD challenge by using results of extensive experiments over two visual relation benchmarks.

READ FULL TEXT

page 7

page 8

research
11/09/2015

Weakly Supervised Deep Detection Networks

Weakly supervised learning of object detection is an important problem i...
research
07/29/2017

Weakly-supervised learning of visual relations

This paper introduces a novel approach for modeling visual relations bet...
research
06/16/2020

Explanation-based Weakly-supervised Learning of Visual Relations with Graph Networks

Visual relationship detection is fundamental for holistic image understa...
research
08/07/2017

Two-Phase Learning for Weakly Supervised Object Localization

Weakly supervised semantic segmentation and localiza- tion have a proble...
research
12/05/2017

R-FCN-3000 at 30fps: Decoupling Detection and Classification

We present R-FCN-3000, a large-scale real-time object detector in which ...
research
08/17/2021

Fully Convolutional Networks for Panoptic Segmentation with Point-based Supervision

In this paper, we present a conceptually simple, strong, and efficient f...
research
04/05/2017

Weakly Supervised Dense Video Captioning

This paper focuses on a novel and challenging vision task, dense video c...

Please sign up or login with your details

Forgot password? Click here to reset