Talk2Car: Taking Control of Your Self-Driving Car

09/24/2019
by   Thierry Deruyttere, et al.
20

A long-term goal of artificial intelligence is to have an agent execute commands communicated through natural language. In many cases the commands are grounded in a visual environment shared by the human who gives the command and the agent. Execution of the command then requires mapping the command into the physical visual space, after which the appropriate action can be taken. In this paper we consider the former. Or more specifically, we consider the problem in an autonomous driving setting, where a passenger requests an action that can be associated with an object found in a street scene. Our work presents the Talk2Car dataset, which is the first object referral dataset that contains commands written in natural language for self-driving cars. We provide a detailed comparison with related datasets such as ReferIt, RefCOCO, RefCOCO+, RefCOCOg, Cityscape-Ref and CLEVR-Ref. Additionally, we include a performance analysis using strong state-of-the-art models. The results show that the proposed object referral task is a challenging one for which the models show promising results but still require additional research in natural language processing, computer vision and the intersection of these fields. The dataset can be found on our website: http://macchina-ai.eu/

READ FULL TEXT

page 2

page 5

page 14

research
12/10/2021

Predicting Physical World Destinations for Commands Given to Self-Driving Cars

In recent years, we have seen significant steps taken in the development...
research
09/18/2020

Commands 4 Autonomous Vehicles (C4AV) Workshop Summary

The task of visual grounding requires locating the most relevant region ...
research
04/26/2022

Landing AI on Networks: An equipment vendor viewpoint on Autonomous Driving Networks

The tremendous achievements of Artificial Intelligence (AI) in computer ...
research
05/28/2017

Listen, Interact and Talk: Learning to Speak via Interaction

One of the long-term goals of artificial intelligence is to build an age...
research
10/16/2019

Conditional Driving from Natural Language Instructions

Widespread adoption of self-driving cars will depend not only on their s...
research
11/21/2017

Mastering the Dungeon: Grounded Language Learning by Mechanical Turker Descent

Contrary to most natural language processing research, which makes use o...
research
10/18/2021

NYU-VPR: Long-Term Visual Place Recognition Benchmark with View Direction and Data Anonymization Influences

Visual place recognition (VPR) is critical in not only localization and ...

Please sign up or login with your details

Forgot password? Click here to reset