Recurrent Vision Transformers for Object Detection with Event Cameras

12/11/2022
by   Mathias Gehrig, et al.
0

We present Recurrent Vision Transformers (RVTs), a novel backbone for object detection with event cameras. Event cameras provide visual information with sub-millisecond latency at a high-dynamic range and with strong robustness against motion blur. These unique properties offer great potential for low-latency object detection and tracking in time-critical scenarios. Prior work in event-based vision has achieved outstanding detection performance but at the cost of substantial inference time, typically beyond 40 milliseconds. By revisiting the high-level design of recurrent vision backbones, we reduce inference time by a factor of 5 while retaining similar performance. To achieve this, we explore a multi-stage design that utilizes three key concepts in each stage: First, a convolutional prior that can be regarded as a conditional positional embedding. Second, local- and dilated global self-attention for spatial feature interaction. Third, recurrent temporal feature aggregation to minimize latency while retaining temporal information. RVTs can be trained from scratch to reach state-of-the-art performance on event-based object detection - achieving an mAP of 47.5 RVTs offer fast inference (13 ms on a T4 GPU) and favorable parameter efficiency (5 times fewer than prior art). Our study brings new insights into effective design choices that could be fruitful for research beyond event-based vision.

READ FULL TEXT

page 8

page 10

research
09/06/2021

Moving Object Detection for Event-based Vision using k-means Clustering

Moving object detection is a crucial task in computer vision. Event-base...
research
09/30/2021

Moving Object Detection for Event-based vision using Graph Spectral Clustering

Moving object detection has been a central topic of discussion in comput...
research
12/15/2022

Event-based Visual Tracking in Dynamic Environments

Visual object tracking under challenging conditions of motion and light ...
research
07/26/2023

Memory-Efficient Graph Convolutional Networks for Object Classification and Detection with Event Cameras

Recent advances in event camera research emphasize processing data in it...
research
12/06/2022

Event-based Monocular Dense Depth Estimation with Recurrent Transformers

Event cameras, offering high temporal resolutions and high dynamic range...
research
12/29/2022

High-temporal-resolution event-based vehicle detection and tracking

Event-based vision has been rapidly growing in recent years justified by...
research
11/22/2022

Pushing the Limits of Asynchronous Graph-based Object Detection with Event Cameras

State-of-the-art machine-learning methods for event cameras treat events...

Please sign up or login with your details

Forgot password? Click here to reset