arrow
Return

Assessing deep learning models for multi-class upper endoscopic disease segmentation: A comprehensive comparative study

delete2025-11-07
delete0
PRE
AI
Y
Yao, Xuming
P
Pak Kin Wong *
晏涛 cover
晏涛 (Tao Yan)
Y
Yanyan Hu
C
Chon In Chan
Y
Ye-Ying Qin
C
Chi Hong Wong
I
In Weng Chan
I
Ieng Hou Lam
W
Wong, Sio Hou
Z
Zheng Li
S
Shan Gao
H
Hon Ho Yu
L
Liang Yao
B
Baoliang Zhao
Y
Ying Hu
DOI:10.3748/wjg.v31.i41.111184delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
BACKGROUND Upper gastrointestinal (UGI) diseases present diagnostic challenges during endoscopy due to visual similarities, indistinct boundaries, and observer variability, which can lead to missed diagnoses and delayed treatment. Automated segmentation using deep learning (DL) models offers the potential to assist endoscopists, improve diagnostic accuracy, and reduce workload. However, multi-class UGI disease segmentation remains underexplored, with limited annotated datasets and insufficient focus on clinical validation. This study hypothesizes that comparative analysis of different DL architectures can identify models suitable for clinical application, providing actionable insights to reduce diagnostic errors and support clinical decision-making in endoscopic practice. AIM To evaluate 17 state-of-the-art DL models for multi-class UGI disease segmentation, emphasizing clinical translation and real-world applicability.METHODS This study evaluated 17 DL models spanning convolutional neural network (CNN)-, transformer-, and mamba-based architectures using a self-collected dataset from two hospitals in Macao and Xiangyang (3313 images, 9 classes) and the public EDD2020 dataset (386 images, 5 classes). Models were assessed for segmentation performance and performance-efficiency trade-off. Statistical analyses were conducted to examine performance differences across architectures. Generalization capability was measured through a cross-dataset evaluation (training models on the self-collected dataset and testing on the EDD2020 dataset). RESULTS Swin-UMamba achieved the highest segmentation performance across both datasets [intersection over union (IoU): 89.06% +/- 0.20% self-collected, 77.53% +/- 0.32% EDD2020], followed by SegFormer (IoU: 88.94% +/- 0.38% self-collected, 77.20% +/- 0.98% EDD2020) and ConvNeXt + UPerNet (IoU: 88.48% +/- 0.09% self-collected, 76.90% +/- 0.61% EDD2020). Statistical analyses showed no significant differences between paradigms, though hierarchical architectures with pre-trained encoders consistently outperformed simpler designs. SegFormer achieved the best balance of accuracy and computational efficiency with a performance-efficiency trade-off score of 92.02%, making it suitable for real-time clinical use. Cross-dataset evaluation revealed significant performance drops, with generalization retention rates of 64.78% to 71.52%. Transformer-based models, particularly pyramid vision transformer v2 + efficient multi-scale convolutional decoding (IoU: 63.35% +/- 1.44%), generalized better than CNN- and mamba-based models. CONCLUSION Hierarchical architectures like Swin-UMamba and SegFormer show promise for UGI disease segmentation, reducing missed diagnoses and improving workflows, but robust clinical validation is crucial for real-world deployment.
Keywords:
Deep learning
Upper endoscopy
Medical imaging
Gastrointestinal diseases
Disease segmentation

Journal

World Journal of Gastroenterology cover
World Journal of Gastroenterology
IF:
5.4
Papers:
2.1W
Citations:
5.1W

Organization

H
hubei university of arts & science
Scholars:
1.9K
Papers: 1.5K
Citations: 3
S
shenzhen institute of advanced technology, cas
Scholars:
5.6K
Papers: 4.5K
Citations: 7
U
university of macau
Scholars:
2.5K
Papers: 1.4K
Citations: 0
C
chinese academy of sciences
Scholars:
56.5W
Papers: 44.9W
Citations: 704
researcher View more organizations