Time3D: End-to-End Joint Monocular 3D Object Detection and Tracking for Autonomous Driving

05/30/2022
by   Peixuan Li, et al.
0

While separately leveraging monocular 3D object detection and 2D multi-object tracking can be straightforwardly applied to sequence images in a frame-by-frame fashion, stand-alone tracker cuts off the transmission of the uncertainty from the 3D detector to tracking while cannot pass tracking error differentials back to the 3D detector. In this work, we propose jointly training 3D detection and 3D tracking from only monocular videos in an end-to-end manner. The key component is a novel spatial-temporal information flow module that aggregates geometric and appearance features to predict robust similarity scores across all objects in current and past frames. Specifically, we leverage the attention mechanism of the transformer, in which self-attention aggregates the spatial information in a specific frame, and cross-attention exploits relation and affinities of all objects in the temporal domain of sequence frames. The affinities are then supervised to estimate the trajectory and guide the flow of information between corresponding 3D objects. In addition, we propose a temporal -consistency loss that explicitly involves 3D target motion modeling into the learning, making the 3D trajectory smooth in the world coordinate system. Time3D achieves 21.4% AMOTA, 13.6% AMOTP on the nuScenes 3D tracking benchmark, surpassing all published competitors, and running at 38 FPS, while Time3D achieves 31.2% mAP, 39.4% NDS on the nuScenes 3D detection benchmark.

READ FULL TEXT

page 1

page 5

page 8

research
11/27/2020

Temporal-Channel Transformer for 3D Lidar-Based Video Object Detection in Autonomous Driving

The strong demand of autonomous driving in the industry has lead to stro...
research
11/03/2017

End-to-end Flow Correlation Tracking with Spatial-temporal Attention

Discriminative correlation filters (DCF) with deep convolutional feature...
research
04/02/2020

Tracking Objects as Points

Tracking has traditionally been the art of following interest points thr...
research
08/09/2017

Online Multi-Object Tracking Using CNN-based Single Object Tracker with Spatial-Temporal Attention Mechanism

In this paper, we propose a CNN-based framework for online MOT. This fra...
research
06/09/2023

TrajectoryFormer: 3D Object Tracking Transformer with Predictive Trajectory Hypotheses

3D multi-object tracking (MOT) is vital for many applications including ...
research
07/12/2023

Multi-Object Tracking as Attention Mechanism

We propose a conceptually simple and thus fast multi-object tracking (MO...
research
01/06/2018

ReMotENet: Efficient Relevant Motion Event Detection for Large-scale Home Surveillance Videos

This paper addresses the problem of detecting relevant motion caused by ...

Please sign up or login with your details

Forgot password? Click here to reset