arrow
返回

RT-ViT: Real-Time Monocular Depth Estimation Using Lightweight Vision Transformers

delete2022-05-19
delete12
delete
OA
AI
H
Hatem Ibrahem
A
Ahmed Salem
H
Hyun‐Soo Kang *
DOI:10.3390/s22103849delete
delete原文链接
delete分享
delete收藏
查看原文
摘要

摘要

En 中文
The latest research in computer vision highlighted the effectiveness of the vision transformers (ViT) in performing several computer vision tasks; they can efficiently understand and process the image globally unlike the convolution which processes the image locally. ViTs outperform the convolutional neural networks in terms of accuracy in many computer vision tasks but the speed of ViTs is still an issue, due to the excessive use of the transformer layers that include many fully connected layers. Therefore, we propose a real-time ViT-based monocular depth estimation (depth estimation from single RGB image) method with encoder-decoder architectures for indoor and outdoor scenes. This main architecture of the proposed method consists of a vision transformer encoder and a convolutional neural network decoder. We started by training the base vision transformer (ViT-b16) with 12 transformer layers then we reduced the transformer layers to six layers, namely ViT-s16 (the Small ViT) and four layers, namely ViT-t16 (the Tiny ViT) to obtain real-time processing. We also try four different configurations of the CNN decoder network. The proposed architectures can learn the task of depth estimation efficiently and can produce more accurate depth predictions than the fully convolutional-based methods taking advantage of the multi-head self-attention module. We train the proposed encoder-decoder architecture end-to-end on the challenging NYU-depthV2 and CITYSCAPES benchmarks then we evaluate the trained models on the validation and test sets of the same benchmarks showing that it outperforms many state-of-the-art methods on depth estimation while performing the task in real-time (similar to 20 fps). We also present a fast 3D reconstruction (similar to 17 fps) experiment based on the depth estimated from our method which is considered a real-world application of our method.
Keyword:
monocular depth estimation
convolutional neural networks
vision transformers
real-time processing
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

Sensors 封面图
Sensors
IF:
3.5
论文数:
7.2W
被引数:
20.9W

机构

C
Chungbuk National University
学者数:
8.5K
论文数: 8.0K
被引数: 6.4K
引用论文

引用论文

Psychometric properties of a structured interview guide for the rating for anxiety in dementia
err2012-02-28
err0
errOAAI
errA. Lynn Snow; Cashuna Huddleston; Christina Robinson; Mark E. Kunik; Amber L. Bush; Nancy Wilson; Jessica Calleo; Amber Paukert; Cynthia Kraus-Schuman; Nancy J. Petersen; Melinda A. Stanley
err分享
err收藏
3D Ken Burns Effect from a Single Image
err2019-11-08
err137
errOAAI
errNiklaus, Simon; Mai, Long; Yang, Jimei; Liu, Feng
err分享
err收藏
Effects of Machining Parameters on the Quality in Machining of Aluminium Alloys Thin Plates
err2019-08-24
err0
errOAAI
errIrene Del Sol; Asuncion Rivero; Antonio J. Gamez
err分享
err收藏
Lifestyle behaviors and associated factors among individuals with diabetes in Brazil: a latent class analysis approach
err2023-07-01
err0
errOAAI
errGabriela Bertoldi Peres; Luciana Bertoldi Nucci; André Luiz Monezi Andrade; Carla Cristina Enes
err分享
err收藏
Ontop-spatial: Ontop of geospatial databases
err2019-10-01
err0
PREAI
errKonstantina Bereta; Guohui Xiao; Manolis Koubarakis
err分享
err收藏
ImageNet Large Scale Visual Recognition ChallengeImageNet大规模视觉识别挑战
err2015-04-11
err2.7W
PREAI
errRussakovsky, Olga; Deng, Jia; Su, Hao; Krause, Jonathan; Satheesh, Sanjeev; Ma, Sean; Huang, Zhiheng; Karpathy, Andrej; Khosla, Aditya; Bernstein, Michael; Berg, Alexander C.; Fei-Fei, Li
err分享
err收藏
Fire Detection Method in Smart City Environments Using a Deep-Learning-Based Approach
err2021-12-27
err0
errOAAI
errKuldoshbay Avazov; Mukhriddin Mukhiddinov; Fazliddin Makhmudov; Young Im Cho
err分享
err收藏
学者 查看更多内容