Return
TMSA-Net: Transformer-Based Multi-Scale Attention U-Net for Flood Image Segmentation
P
P
A
C
DOI:10.1049/cit2.70169.png)
Abstract
En 中文
Flood detection is essential for real-time applications, including disaster management, emergency response, and alerting people in flood zones. For successful flood detection, accurate flood region segmentation is essential. However, the flood region segmentation is challenging due to the complex background and occlusions with debris and the effect of external adverse factors. In contrast to existing models that focus on satellite and remote sensing images, for real-time applications, this study proposes a new Transformer-based Multi-Scale-Attention U-Net (TMSA-Net) model for flood region segmentation in images with cluttered backgrounds. The proposed model extracts features hierarchically using a Convolutional Neural Network (CNN) encoder. The deepest features are fed to an adapted vision transformer for extracting global context. In the decoder, Multi-Scale Attention Gates are used in skip connections to selectively filter encoder features before fusion. The proposed model is compared with state-of-the-art methods, including Swin-UNet, DeepLabV3+, TransUNet, UNet++, and Attention U-Net. Experimental results show that TMSA-Net achieves a test IoU of 92.01%, with improvements of 2.82% and 5.95% IoU over DeepLabV3+ and TransUNet, respectively, demonstrating the effectiveness of the proposed approach. The newly created experimental dataset, with ground truth at the pixel level, will be publicly available (released upon acceptance).
Keywords:
attention mechanism
deep learning
disaster response
flood detection
multi-scale processing
semantic segmentation
vision transformer
AI Summary
Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.
Journal
IF:
7.3
Papers:
649
Citations:
2.4K
