arrow
Return

Bidirectional interactive multi-scale network using Wave-Conv ViT for single image deraining

delete2025-04-01
delete0
PRE
AI
S
Siyan Fang *
刘斌 (Bin Liu)
DOI:10.1016/j.image.2025.117311delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
To address the limitations of high-frequency information capture by Vision Transformer (ViT) and the loss of fine details in existing image deraining methods, we introduce a Bidirectional Interactive Multi-Scale Network (BIMNet) that employs newly developed Wave-Conv ViT (WCV). The WCV utilizes a wavelet transform to enable self-attention in both low-frequency and high-frequency domains, significantly enhancing ViT's capacity for diverse frequency-domain feature modeling. Additionally, by incorporating convolutional operations, WCV enhances the extraction and integration of local features across various spatial windows. BIMNet injects rainy images into deep network layers, enabling bidirectional propagation with shallow layer features that enrich skip connections with detailed and complementary information, thus improving the fidelity of detail recovery. Moreover, we present the CORain1000 dataset, tailored for the dual challenges of image deraining and object detection, which offers more diversity in rain patterns, image sizes, and volumes than the commonly used COCO350 dataset. Extensive experiments demonstrate the superiority of BIMNet over advanced methods. The code and CORain1000 dataset are available at https://github.com/fashyon/BIMNet.
Keywords:
Image deraining
Vision Transformer
Wavelet
Convolution
Multi-scale

Journal

S
Signal Processing and Image Communication
IF:
2.7
Papers:
2.8K
Citations:
4.2K

Organization

H
hubei university
Scholars:
1.1W
Papers: 7.0K
Citations: 7