1
Return

Frequency-based cross-attention fusion network for RGB-D salient object detection

delete2026-07-01
delete0
PRE
AI
X
Xin Zhou *
W
Wenyao Ji
DOI:10.1016/j.neunet.2026.109322delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
In RGB-D salient object detection tasks, effectively fusing cross-modality information is crucial for improving detection performance. Current cross-modality fusion methods primarily focus on integrating the RGB and depth feature contents in the spatial domain, while the exploration of these fusion methods in the frequency domain remains relatively limited. To address this challenge, this paper proposes a Frequency-based Cross-Attention Fusion Network for RGB-D salient object detection (FCAFNet). Specifically, (1) a Frequency-based element-product Cross-Attention Module (FCAM) is employed to establish the long-range feature relationship between tokens from RGB and depth modality. (2) A Bi-directional Feature Aggregation Module (BFAM) is proposed to aggregate the horizontal and vertical feature contexts using different-size convolutional kernels. Finally, to refine the boundaries of salient objects, an edge supervision module (ESM) is proposed to sharpen the edge contour by the application of high- and low-frequency features in the bottom layers. Experimental results demonstrate that the proposed model outperforms the state-of-the-art models qualitatively and quantitatively. Our code is publicly available at: https://github.com/zhx-hub/FCAFNet
Keywords:
RGB-D
Salient object detection
Frequency domain
Cross-attention fusion
Edge supervision
Bi-directional feature aggregation

Journal

Neural Networks cover
Neural Networks
IF:
6.3
Papers:
7.7K
Citations:
3.0W

Organization

U
University of Cologne
Scholars:
2.9W
Papers: 2.1W
Citations: 2.4W
W
western university
Scholars:
1.1K
Papers: 576
Citations: 0
Cited Papers

Cited Papers

Citing Papers

Citing Papers