Log In Sign Up

SIRI: Spatial Relation Induced Network For Spatial Description Resolution

by   Peiyao Wang, et al.

Spatial Description Resolution, as a language-guided localization task, is proposed for target location in a panoramic street view, given corresponding language descriptions. Explicitly characterizing an object-level relationship while distilling spatial relationships are currently absent but crucial to this task. Mimicking humans, who sequentially traverse spatial relationship words and objects with a first-person view to locate their target, we propose a novel spatial relationship induced (SIRI) network. Specifically, visual features are firstly correlated at an implicit object-level in a projected latent space; then they are distilled by each spatial relationship word, resulting in each differently activated feature representing each spatial relationship. Further, we introduce global position priors to fix the absence of positional information, which may result in global positional reasoning ambiguities. Both the linguistic and visual features are concatenated to finalize the target localization. Experimental results on the Touchdown show that our method is around 24% better than the state-of-the-art method in terms of accuracy, measured by an 80-pixel radius. Our method also generalizes well on our proposed extended dataset collected using the same settings as Touchdown.


page 2

page 4

page 6

page 7

page 12

page 13

page 14

page 15


Exploring the Semantics for Visual Relationship Detection

Scene graph construction / visual relationship detection from an image a...

Natural Language Guided Visual Relationship Detection

Reasoning about the relationships between object pairs in images is a cr...

Interpretable Visual Reasoning via Induced Symbolic Space

We study the problem of concept induction in visual reasoning, i.e., ide...

On Exploring Undetermined Relationships for Visual Relationship Detection

In visual relationship detection, human-notated relationships can be reg...

Improving Robot Localisation by Ignoring Visual Distraction

Attention is an important component of modern deep learning. However, le...

Acquiring Common Sense Spatial Knowledge through Implicit Spatial Templates

Spatial understanding is a fundamental problem with wide-reaching real-w...

Semantic Signatures for Large-scale Visual Localization

Visual localization is a useful alternative to standard localization tec...