arrow
返回

Self-Supervised Monocular Depth Estimation Using Hybrid Transformer Encoder

delete2022-10-01
delete16
delete
OA
AI
S
Seung-Jun Hwang
S
Sung Jun Park
J
Joong-Hwan Baek *
B
Byungkyu Kim
DOI:10.1109/JSEN.2022.3199265delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Depth estimation using monocular camera sensors is an important technique in computer vision. Supervised monocular depth estimation requires a lot of data acquired from depth sensors. However, acquiring depth data is an expensive task. We sometimes cannot acquire data due to the limitations of the sensor. View synthesis-based depth estimation research is a self-supervised learning method that does not require depth data supervision. Previous studies mainly use the convolutional neural network (CNN)-based networks in encoders. The CNN is suitable for extracting local features through convolution operation. Recent vision transformers (ViTs) are suitable for global feature extraction based on multiself-attention modules. In this article, we propose a hybrid network combining the CNN and ViT networks in self-supervised learning-based monocular depth estimation. We design an encoder-decoder structure that uses CNNs in the earlier stage of extracting local features and a ViT in the later stages of extracting global features. We evaluate the proposed network through various experiments based on the Karlsruhe Institute of Technology and Toyota Technological Institute (KITTI) and Cityscapes datasets. The results showed higher performance than previous studies and reduced parameters and computations. Codes and trained models are available at https://github.com/fogfog2/manydepthformer.
Keyword:
Estimation
Transformers
Feature extraction
Computational modeling
Cameras
Image reconstruction
Costs
Depth estimation
monocular sensor estimation
self-attention
self-supervised
transformer

期刊

IEEE Sensors Journal 封面图
IEEE Sensors Journal
IF:
4.5
论文数:
2.1W
被引数:
7.3W

机构

K
Korea Aerospace University
学者数:
1.1K
论文数: 1.1K
被引数: 513