Return
EVSSD: Efficient visual state space decoding for 2D medical image segmentation
DOI:10.1016/j.sigpro.2026.110760.png)
Abstract
En 中文
Transformers have gained widespread attention in medical image segmentation due to the self-attention, which effectively captures global contextual information. However, self-attention requires pairwise interactions among all pixels, leading to quadratic computational complexity. This results in substantial computational overhead when processing high-resolution medical images, posing a significant challenge to the efficient segmentation in clinical applications. To address this issue, we propose an Efficient Visual State Space Decoding, called EVSSD. EVSSD comprises three key components: the Efficient Upsampling Module (EUM), the Efficient Mamba Block (EMB), and the Guided Fusion Module (GFM). EUM could achieve efficient upsampling of decoded features with fewer parameters through pixel shuffle along the channel dimension. The EMB employs a visual state space model with linear complexity to capture global contextual information in images. While the GFM adaptively leverages decoded features to guide and fuse encoded features, effectively bridging the gap between low-level details and high-level semantics. Moreover, we employ a multi-stage aggregation loss to optimize the model. Experimental results show that EVSSD achieves state-of-the-art performance on multimodal (CT, endoscopic images, dermoscopic images) image datasets such as Synapse, Kvasir-Seg and ISIC2018. For reproducibility, the source code is available at https://github.com/XYQ1517/EVSSD .
Keywords:
Medical image segmentation
Efficient decoding
Visual state space model
Multi-stage aggregation loss
Multi-modal image datasets
Journal
IF:
3.6
Papers:
9.9K
Citations:
1.7W

