arrow
Return

Query-Driven Retinal Layer Segmentation in OCT Using Cross-Attentive Feature Learning

delete2026-05-31
delete0
delete
OA
AI
N
Nebras Sobahi
S
Salih Taha Alperen Özçelik *
O
Orhan Atıla
A
Abdulkadir Şengür
M
Muhammed Halil Akpınar
DOI:10.3390/diagnostics16111697delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Background/Objectives: Retinal layer segmentation in optical coherence tomography (OCT) is essential for the diagnosis and monitoring of retinal diseases such as age-related macular degeneration (AMD) and diabetic macular edema (DME). Although deep learning methods have achieved strong performance, most rely on dense pixel-wise predictions and often struggle to preserve anatomical consistency, particularly in regions with low contrast or structural deformation. This study aims to address these limitations by introducing a query-based segmentation framework that explicitly models retinal layer structure. Methods: In this paper, we propose the RetiQueryNet architecture that employs encoding of retinal layers in the form of query embeddings with the use of cross attention to interact with pixel level features encoded by a transformer based encoder. The architecture integrates multi-scale features through a compact query-driven decoder with modest additional computational overhead. Normalization and resizing of OCT images preceded their usage as inputs, while the layer labels were converted to multi-class segmentation maps. In the training process, we used loss function with combination of cross entropy loss and Dice loss. Our model performance was compared with multiple state-of-the-art models such as U-Net, DeepLabV3, FPN, MANet and SegFormer, while performance metrics were Dice, IoU and mean surface distance (MSD). Results: RetiQueryNet was able to attain a mean Dice score of 0.934 ± 0.0046 and outperformed all baseline models on the main performance measures. Improvements were particularly evident in challenging retinal layers such as IBRPE and OBRPE, where boundary ambiguity is high. It should be noted that RetiQueryNet had a relatively lower MSD value, meaning that the predicted boundaries were more accurate. Furthermore, visual observations suggest that the approach generated smooth and coherent segmentations. Conclusions: The findings demonstrate that query-based modeling offers a viable approach to pixel-wise segmentation. In particular, by making use of structural priors in the form of learnable queries, RetiQueryNet improves not only segmentation accuracy but also anatomical consistency. Query-based modeling appears to be an exciting area for retinal image segmentation that could potentially be applied to other applications in medical image segmentation.
Keywords:
OCT
retinal layer segmentation
transformer
query-based learning
cross-attention
medical image segmentation

Journal

Diagnostics cover
Diagnostics
IF:
3.3
Papers:
1.9W
Citations:
3.6W

Organization

F
Firat University
Scholars:
4.1K
Papers: 3.9K
Citations: 43
K
King Abdulaziz University
Scholars:
1.9W
Papers: 1.9W
Citations: 3.3W
B
Bingol University
Scholars:
686
Papers: 940
Citations: 19
I
istanbul university-cerrahpasa
Scholars:
327
Papers: 159
Citations: 0
researcher View more organizations