RandomRooms: Unsupervised Pre-training from Synthetic Shapes and Randomized Layouts for 3D Object Detection

08/17/2021
by   Yongming Rao, et al.
12

3D point cloud understanding has made great progress in recent years. However, one major bottleneck is the scarcity of annotated real datasets, especially compared to 2D object detection tasks, since a large amount of labor is involved in annotating the real scans of a scene. A promising solution to this problem is to make better use of the synthetic dataset, which consists of CAD object models, to boost the learning on real datasets. This can be achieved by the pre-training and fine-tuning procedure. However, recent work on 3D pre-training exhibits failure when transfer features learned on synthetic objects to other real-world applications. In this work, we put forward a new method called RandomRooms to accomplish this objective. In particular, we propose to generate random layouts of a scene by making use of the objects in the synthetic CAD dataset and learn the 3D scene representation by applying object-level contrastive learning on two random scenes generated from the same set of synthetic objects. The model pre-trained in this way can serve as a better initialization when later fine-tuning on the 3D object detection task. Empirically, we show consistent improvement in downstream 3D detection tasks on several base models, especially when less training data are used, which strongly demonstrates the effectiveness and generalization of our method. Benefiting from the rich semantic knowledge and diverse objects from synthetic data, our method establishes the new state-of-the-art on widely-used 3D detection benchmarks ScanNetV2 and SUN RGB-D. We expect our attempt to provide a new perspective for bridging object and scene-level 3D understanding.

READ FULL TEXT

page 11

page 12

research
06/07/2023

Randomized 3D Scene Generation for Generalizable Self-supervised Pre-training

Capturing and labeling real-world 3D data is laborious and time-consumin...
research
07/21/2020

PointContrast: Unsupervised Pre-training for 3D Point Cloud Understanding

Arguably one of the top success stories of deep learning is transfer lea...
research
05/10/2022

UNITS: Unsupervised Intermediate Training Stage for Scene Text Detection

Recent scene text detection methods are almost based on deep learning an...
research
11/21/2018

Rethinking ImageNet Pre-training

We report competitive results on object detection and instance segmentat...
research
12/06/2021

4DContrast: Contrastive Learning with Dynamic Correspondences for 3D Scene Understanding

We present a new approach to instill 4D dynamic object priors into learn...
research
08/14/2023

UniWorld: Autonomous Driving Pre-training via World Models

In this paper, we draw inspiration from Alberto Elfes' pioneering work i...
research
01/02/2021

VinVL: Making Visual Representations Matter in Vision-Language Models

This paper presents a detailed study of improving visual representations...

Please sign up or login with your details

Forgot password? Click here to reset