Weakly Supervised PatchNets: Describing and Aggregating Local Patches for Scene Recognition

09/01/2016
by   Zhe Wang, et al.
0

Traditional feature encoding scheme (e.g., Fisher vector) with local descriptors (e.g., SIFT) and recent convolutional neural networks (CNNs) are two classes of successful methods for image recognition. In this paper, we propose a hybrid representation, which leverages the discriminative capacity of CNNs and the simplicity of descriptor encoding schema for image recognition, with a focus on scene recognition. To this end, we make three main contributions from the following aspects. First, we propose a patch-level and end-to-end architecture to model the appearance of local patches, called PatchNet. PatchNet is essentially a customized network trained in a weakly supervised manner, which uses the image-level supervision to guide the patch-level feature extraction. Second, we present a hybrid visual representation, called VSAD, by utilizing the robust feature representations of PatchNet to describe local patches and exploiting the semantic probabilities of PatchNet to aggregate these local patches into a global representation. Third, based on the proposed VSAD representation, we propose a new state-of-the-art scene recognition approach, which achieves an excellent performance on two standard benchmarks: MIT Indoor67 (86.2%) and SUN397 (73.0%).

READ FULL TEXT

page 5

page 6

page 7

page 12

research
11/29/2016

Weakly-supervised Discriminative Patch Learning via CNN for Fine-grained Recognition

Research on fine-grained recognition has recently shifted from multistag...
research
05/06/2017

Deep Patch Learning for Weakly Supervised Object Classification and Discovery

Patch-level image representation is very important for object classifica...
research
09/16/2022

Weakly Supervised Semantic Segmentation via Progressive Patch Learning

Most of the existing semantic segmentation approaches with image-level c...
research
02/11/2022

Patch-NetVLAD+: Learned patch descriptor and weighted matching strategy for place recognition

Visual Place Recognition (VPR) in areas with similar scenes such as urba...
research
02/21/2017

Scene Recognition by Combining Local and Global Image Descriptors

Object recognition is an important problem in computer vision, having di...
research
10/06/2015

Harvesting Discriminative Meta Objects with Deep CNN Features for Scene Classification

Recent work on scene classification still makes use of generic CNN featu...
research
07/08/2014

Orientation covariant aggregation of local descriptors with embeddings

Image search systems based on local descriptors typically achieve orient...

Please sign up or login with your details

Forgot password? Click here to reset