Cross-View Policy Learning for Street Navigation

06/13/2019
by   Ang Li, et al.
5

The ability to navigate from visual observations in unfamiliar environments is a core component of intelligent agents and an ongoing challenge for Deep Reinforcement Learning (RL). Street View can be a sensible testbed for such RL agents, because it provides real-world photographic imagery at ground level, with diverse street appearances; it has been made into an interactive environment called StreetLearn and used for research on navigation. However, goal-driven street navigation agents have not so far been able to transfer to unseen areas without extensive retraining, and relying on simulation is not a scalable solution. Since aerial images are easily and globally accessible, we propose instead to train a multi-modal policy on ground and aerial views, then transfer the ground view policy to unseen (target) parts of the city by utilizing aerial view observations. Our core idea is to pair the ground view with an aerial view and to learn a joint policy that is transferable across views. We achieve this by learning a similar embedding space for both views, distilling the policy across views and dropping out visual modalities. We further reformulate the transfer learning paradigm into three stages: 1) cross-modal training, when the agent is initially trained on multiple city regions, 2) aerial view-only adaptation to a new area, when the agent is adapted to a held-out region using only the easily obtainable aerial view, and 3) ground view-only transfer, when the agent is tested on navigation tasks on unseen ground views, without aerial imagery. Experimental results suggest that the proposed cross-view policy learning enables better generalization of the agent and allows for more effective transfer to unseen environments.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
07/12/2023

VELMA: Verbalization Embodiment of LLM Agents for Vision and Language Navigation in Street View

Incremental decision making in real-world environments is one of the mos...
research
10/29/2019

Navigation Agents for the Visually Impaired: A Sidewalk Simulator and Experiments

Millions of blind and visually-impaired (BVI) people navigate urban envi...
research
03/01/2019

Learning To Follow Directions in Street View

Navigating and understanding the real world remains a key challenge in m...
research
04/09/2021

Uncovering commercial activity in informal cities

Knowledge of the spatial organisation of economic activity within a city...
research
07/03/2023

MoVie: Visual Model-Based Policy Adaptation for View Generalization

Visual Reinforcement Learning (RL) agents trained on limited views face ...
research
01/10/2023

Pix2Map: Cross-modal Retrieval for Inferring Street Maps from Images

Self-driving vehicles rely on urban street maps for autonomous navigatio...
research
05/10/2022

VesNet-RL: Simulation-based Reinforcement Learning for Real-World US Probe Navigation

Ultrasound (US) is one of the most common medical imaging modalities sin...

Please sign up or login with your details

Forgot password? Click here to reset