MonoViT: Self-Supervised Monocular Depth Estimation with a Vision Transformer

08/06/2022
by   Chaoqiang Zhao, et al.
31

Self-supervised monocular depth estimation is an attractive solution that does not require hard-to-source depth labels for training. Convolutional neural networks (CNNs) have recently achieved great success in this task. However, their limited receptive field constrains existing network architectures to reason only locally, dampening the effectiveness of the self-supervised paradigm. In the light of the recent successes achieved by Vision Transformers (ViTs), we propose MonoViT, a brand-new framework combining the global reasoning enabled by ViT models with the flexibility of self-supervised monocular depth estimation. By combining plain convolutions with Transformer blocks, our model can reason locally and globally, yielding depth prediction at a higher level of detail and accuracy, allowing MonoViT to achieve state-of-the-art performance on the established KITTI dataset. Moreover, MonoViT proves its superior generalization capacities on other datasets such as Make3D and DrivingStereo.

READ FULL TEXT

page 1

page 3

page 4

page 7

page 8

research
02/07/2022

Transformers in Self-Supervised Monocular Depth Estimation with Unknown Camera Intrinsics

The advent of autonomous driving and advanced driver assistance systems ...
research
09/22/2021

Improving 360 Monocular Depth Estimation via Non-local Dense Prediction Transformer and Joint Supervised and Self-supervised Learning

Due to difficulties in acquiring ground truth depth of equirectangular (...
research
11/23/2022

Lite-Mono: A Lightweight CNN and Transformer Architecture for Self-Supervised Monocular Depth Estimation

Self-supervised monocular depth estimation that does not require ground-...
research
02/20/2023

GlocalFuse-Depth: Fusing Transformers and CNNs for All-day Self-supervised Monocular Depth Estimation

In recent years, self-supervised monocular depth estimation has drawn mu...
research
04/28/2022

Depth Estimation with Simplified Transformer

Transformer and its variants have shown state-of-the-art results in many...
research
11/20/2022

Hybrid Transformer Based Feature Fusion for Self-Supervised Monocular Depth Estimation

With an unprecedented increase in the number of agents and systems that ...
research
05/23/2022

MonoFormer: Towards Generalization of self-supervised monocular depth estimation with Transformers

Self-supervised monocular depth estimation has been widely studied recen...

Please sign up or login with your details

Forgot password? Click here to reset