Situational Fusion of Visual Representation for Visual Navigation

08/24/2019
by   William B. Shen, et al.
12

A complex visual navigation task puts an agent in different situations which call for a diverse range of visual perception abilities. For example, to "go to the nearest chair", the agent might need to identify a chair in a living room using semantics, follow along a hallway using vanishing point cues, and avoid obstacles using depth. Therefore, utilizing the appropriate visual perception abilities based on a situational understanding of the visual environment can empower these navigation models in unseen visual environments. We propose to train an agent to fuse a large set of visual representations that correspond to diverse visual perception abilities. To fully utilize each representation, we develop an action-level representation fusion scheme, which predicts an action candidate from each representation and adaptively consolidate these action candidates into the final action. Furthermore, we employ a data-driven inter-task affinity regularization to reduce redundancies and improve generalization. Our approach leads to a significantly improved performance in novel environments over ImageNet-pretrained baseline and other fusion methods.

READ FULL TEXT

page 1

page 4

page 7

research
02/18/2023

VLN-Trans: Translator for the Vision and Language Navigation Agent

Language understanding is essential for the navigation agent to follow i...
research
12/02/2022

Private Multiparty Perception for Navigation

We introduce a framework for navigating through cluttered environments b...
research
07/05/2022

CLEAR: Improving Vision-Language Navigation with Cross-Lingual, Environment-Agnostic Representations

Vision-and-Language Navigation (VLN) tasks require an agent to navigate ...
research
02/02/2022

Image-based Navigation in Real-World Environments via Multiple Mid-level Representations: Fusion Models, Benchmark and Efficient Evaluation

Navigating complex indoor environments requires a deep understanding of ...
research
01/07/2018

Building Generalizable Agents with a Realistic and Rich 3D Environment

Towards bridging the gap between machine and human intelligence, it is o...
research
10/03/2018

Grounding the Experience of a Visual Field through Sensorimotor Contingencies

Artificial perception is traditionally handled by hand-designing task sp...
research
01/30/2021

Enacted Visual Perception: A Computational Model based on Piaget Equilibrium

In Maurice Merleau-Ponty's phenomenology of perception, analysis of perc...

Please sign up or login with your details

Forgot password? Click here to reset