Detecting 11K Classes: Large Scale Object Detection without Fine-Grained Bounding Boxes

08/14/2019
by   Hao Yang, et al.
2

Recent advances in deep learning greatly boost the performance of object detection. State-of-the-art methods such as Faster-RCNN, FPN and R-FCN have achieved high accuracy in challenging benchmark datasets. However, these methods require fully annotated object bounding boxes for training, which are incredibly hard to scale up due to the high annotation cost. Weakly-supervised methods, on the other hand, only require image-level labels for training, but the performance is far below their fully-supervised counterparts. In this paper, we propose a semi-supervised large scale fine-grained detection method, which only needs bounding box annotations of a smaller number of coarse-grained classes and image-level labels of large scale fine-grained classes, and can detect all classes at nearly fully-supervised accuracy. We achieve this by utilizing the correlations between coarse-grained and fine-grained classes with shared backbone, soft-attention based proposal re-ranking, and a dual-level memory module. Experiment results show that our methods can achieve close accuracy on object detection to state-of-the-art fully-supervised methods on two large scale datasets, ImageNet and OpenImages, with only a small fraction of fully annotated classes.

READ FULL TEXT
research
02/28/2017

Weakly- and Semi-Supervised Object Detection with Expectation-Maximization Algorithm

Object detection when provided image-level labels instead of instance-le...
research
03/16/2023

Commonsense Knowledge Assisted Deep Learning for Resource-constrained and Fine-grained Object Detection

In this paper, we consider fine-grained image object detection in resour...
research
12/05/2017

R-FCN-3000 at 30fps: Decoupling Detection and Classification

We present R-FCN-3000, a large-scale real-time object detector in which ...
research
04/16/2012

Large-Scale Automatic Labeling of Video Events with Verbs Based on Event-Participant Interaction

We present an approach to labeling short video clips with English verbs ...
research
05/20/2016

Fine-Grained Classification of Pedestrians in Video: Benchmark and State of the Art

A video dataset that is designed to study fine-grained categorisation of...
research
07/05/2018

Open Logo Detection Challenge

Existing logo detection benchmarks consider artificial deployment scenar...
research
05/25/2019

Efficient Object Annotation via Speaking and Pointing

Deep neural networks deliver state-of-the-art visual recognition, but th...

Please sign up or login with your details

Forgot password? Click here to reset