arrow
返回

Triple-Supervised Convolutional Transformer Aggregation for Robust Monocular Endoscopic Dense Depth Estimation

delete2024-08-01
delete1
PRE
AI
W
Wenkang Fan
W
Wenjing Jiang
H
Hong Shi
H
Huiqing Zeng *
Y
Yinran Chen
X
Xióngbiāo Luó *
DOI:10.1109/TMRB.2024.3407384delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Accurate deeply learned dense depth prediction remains a challenge to monocular vision reconstruction. Compared to monocular depth estimation from natural images, endoscopic dense depth prediction is even more challenging. While it is difficult to annotate endoscopic video data for supervised learning, endoscopic video images certainly suffer from illumination variations (limited lighting source, limited field of viewing, and specular highlight), smooth and textureless surfaces in surgical complex fields. This work explores a new deep learning framework of triple-supervised convolutional transformer aggregation (TSCTA) for monocular endoscopic dense depth recovery without annotating any data. Specifically, TSCTA creates convolutional transformer aggregation networks with a new hybrid encoder that combines dense convolution and scalable transformers to parallel extract local texture features and global spatial-temporal features, while it builds a local and global aggregation decoder to effectively aggregate global features and local features from coarse to fine. Moreover, we develop a self-supervised learning framework with triple supervision, which integrates minimum photometric consistency and depth consistency with sparse depth self-supervision to train our model by unannotated data. We evaluated TSCTA on unannotated monocular endoscopic images collected from various surgical procedures, with the experimental results showing that our methods can achieve more accurate depth range, more complete depth distribution, more sufficient textures, better qualitative and quantitative assessment results than state-of-the-art deeply learned monocular dense depth estimation methods.
Keyword:
Feature extraction
Transformers
Estimation
Convolution
Convolutional codes
Lighting
Unsupervised learning
Monocular depth estimation
vision transformers
self-supervised learning
robotic-assisted endoscopy

期刊

I
IEEE Transactions on Medical Robotics and Bionics
IF:
3.8
论文数:
825
被引数:
1.8K

机构

F
fujian medical university
学者数:
2.9W
论文数: 1.3W
被引数: 13
X
xiamen university
学者数:
5.9W
论文数: 3.8W
被引数: 67
引用论文

引用论文

Evaluation and Stability Analysis of Video-Based Navigation System for Functional Endoscopic Sinus Surgery on In Vivo Clinical Data
err2018-10-01
err55
PREAI
errLeonard, Simon; Sinha, Ayushi; Reiter, Austin; Ishii, Masaru; Gallia, Gary L.; Taylor, Russell H.; Hager, Gregory D.
err分享
err收藏
Evolution of the COPD Assessment Test Score during Chronic Obstructive Pulmonary Disease Exacerbations: Determinants and Prognostic Value
err2013-01-01
err0
errOAAI
errDarwin Feliz-Rodriguez; Santiago Zudaire; Carlos Carpio; Elizabet Martínez; Antonia Gómez-Mendieta; Ana Santiago; Rodolfo Alvarez-Sala; Francisco García-Río
err分享
err收藏
A Concise Synthesis of Globotriaosylsphingosine
err2011-02-11
err0
PREAI
errHenrik Gold; Rolf G. Boot; Johannes M. F. G. Aerts; Herman S. Overkleeft; Jeroen D. C. Codée; Gijs A. van der Marel
err分享
err收藏
err分享
err收藏
err分享
err收藏
err分享
err收藏
err分享
err收藏
学者 查看更多内容