arrow
返回

Multi-level efficient 3D image reconstruction model based on ViT

delete2024-09-04
delete0
delete
OA
AI
R
Ren-Hao Zhang
B
Bingliang Hu *
张耿 (Geng Zhang)
S
Siyuan Li
B
Baocheng Chen
J
Jia Liu
X
Xing Wang
C
Chang Su
X
Xijie Li
张宁 封面图
张宁 (Ning Zhang)
K
Kai Qiao
DOI:10.1364/OE.535211delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Single-photon LIDAR faces challenges in high-quality 3D reconstruction due to high noise levels, low accuracy, and long inference times. Traditional methods, which rely on statistical data to obtain parameter information, are inefficient in high-noise environments. Although convolutional neural networks (CNNs)-based deep learning methods can improve 3D reconstruction quality compared to traditional methods, they struggle to effectively capture global features and long-range dependencies. To address these issues, this paper proposes a multi-level efficient 3D image reconstruction model based on vision transformer (ViT). This model leverages the self-attention mechanism of ViT to capture both global and local features and utilizes attention mechanisms to fuse and refine the extracted features. By introducing generative adversarial ngenerative adversarial networks (GANs), the reconstruction quality and robustness of the model in high noise and low photon environments are further improved. Furthermore, the proposed 3D reconstruction network has been applied in real-world imaging systems, significantly enhancing the imaging capabilities of single-photon 3D reconstruction under strong noise conditions. (c) 2024 Optica Publishing Group under the terms of the Optica Open Access Publishing Agreement

期刊

Optics Express 封面图
Optics Express
IF:
3.3
论文数:
6.1W
被引数:
14.3W

机构

U
university of chinese academy of sciences, cas
学者数:
4.1W
论文数: 3.8W
被引数: 75
X
xi'an institute of optics & precision mechanics, cas
学者数:
536
论文数: 480
被引数: 0
C
chinese academy of sciences
学者数:
56.7W
论文数: 44.9W
被引数: 704
学者 查看更多机构