Structured Visual Search via Composition-aware Learning

10/27/2020
by   Mert Kilickaya, et al.
12

This paper studies visual search using structured queries. The structure is in the form of a 2D composition that encodes the position and the category of the objects. The transformation of the position and the category of the objects leads to a continuous-valued relationship between visual compositions, which carries highly beneficial information, although not leveraged by previous techniques. To that end, in this work, our goal is to leverage these continuous relationships by using the notion of symmetry in equivariance. Our model output is trained to change symmetrically with respect to the input transformations, leading to a sensitive feature space. Doing so leads to a highly efficient search technique, as our approach learns from fewer data using a smaller feature space. Experiments on two large-scale benchmarks of MS-COCO and HICO-DET demonstrates that our approach leads to a considerable gain in the performance against competing techniques.

READ FULL TEXT

page 1

page 8

research
12/10/2020

Tensor Composition Net for Visual Relationship Prediction

We present a novel Tensor Composition Network (TCN) to predict visual re...
research
06/23/2016

Picture It In Your Mind: Generating High Level Visual Representations From Textual Descriptions

In this paper we tackle the problem of image search when the query is a ...
research
04/24/2022

Learning Symmetric Embeddings for Equivariant World Models

Incorporating symmetries can lead to highly data-efficient and generaliz...
research
08/02/2019

Recognizing Image Objects by Relational Analysis Using Heterogeneous Superpixels and Deep Convolutional Features

Superpixel-based methodologies have become increasingly popular in compu...
research
06/15/2021

Compositional Sketch Search

We present an algorithm for searching image collections using free-hand ...
research
12/31/2018

Learning Spatial Common Sense with Geometry-Aware Recurrent Networks

We integrate two powerful ideas, geometry and deep visual representation...
research
10/26/2022

A Sign That Spells: DALL-E 2, Invisual Images and The Racial Politics of Feature Space

In this paper, we examine how generative machine learning systems produc...

Please sign up or login with your details

Forgot password? Click here to reset