SIRI: Spatial Relation Induced Network For Spatial Description Resolution

10/27/2020
by   Peiyao Wang, et al.
0

Spatial Description Resolution, as a language-guided localization task, is proposed for target location in a panoramic street view, given corresponding language descriptions. Explicitly characterizing an object-level relationship while distilling spatial relationships are currently absent but crucial to this task. Mimicking humans, who sequentially traverse spatial relationship words and objects with a first-person view to locate their target, we propose a novel spatial relationship induced (SIRI) network. Specifically, visual features are firstly correlated at an implicit object-level in a projected latent space; then they are distilled by each spatial relationship word, resulting in each differently activated feature representing each spatial relationship. Further, we introduce global position priors to fix the absence of positional information, which may result in global positional reasoning ambiguities. Both the linguistic and visual features are concatenated to finalize the target localization. Experimental results on the Touchdown show that our method is around 24% better than the state-of-the-art method in terms of accuracy, measured by an 80-pixel radius. Our method also generalizes well on our proposed extended dataset collected using the same settings as Touchdown.

READ FULL TEXT

page 2

page 4

page 6

page 7

page 12

page 13

page 14

page 15

research
04/03/2019

Exploring the Semantics for Visual Relationship Detection

Scene graph construction / visual relationship detection from an image a...
research
11/23/2020

Interpretable Visual Reasoning via Induced Symbolic Space

We study the problem of concept induction in visual reasoning, i.e., ide...
research
05/05/2019

On Exploring Undetermined Relationships for Visual Relationship Detection

In visual relationship detection, human-notated relationships can be reg...
research
08/05/2018

Improving Deep Visual Representation for Person Re-identification by Global and Local Image-language Association

Person re-identification is an important task that requires learning dis...
research
09/02/2020

Intrinsic Relationship Reasoning for Small Object Detection

The small objects in images and videos are usually not independent indiv...
research
03/08/2021

Relationship-based Neural Baby Talk

Understanding interactions between objects in an image is an important e...
research
05/07/2020

Semantic Signatures for Large-scale Visual Localization

Visual localization is a useful alternative to standard localization tec...

Please sign up or login with your details

Forgot password? Click here to reset