Return
Cross-Modal Fusion With Mixture-of-Experts for Efficient RGB-D Salient Object Detection
J
F
M
H
DOI:10.1109/tmm.2026.3668533.png)
Abstract
En 中文
Current RGB-D salient object detection (SOD) models are plagued by issues including excessive model parameters and high computational complexity. These drawbacks impede the model’s efficient deployment, and constraining the enhancement of model performance. This paper introduces an efficient and lightweight cross-modal feature cross-fusion network, termed CMFNet. In particular, we design an efficient model utilizing MobileViT as the dual-stream backbone network, thereby significantly reducing computational complexity while maintaining robust feature extraction capabilities. Firstly, we propose a Cross-Fusion Module (CFM) designed to dynamically integrate multi-scale semantic information from RGB and depth features. Furthermore, we design a Lightweight-Mixture-of-Experts Module (L-MoE) for multi-modal features, which enhances the representational capacity of fused features at various levels by employing a dynamic routing mechanism and balanced constraint strategy to allocate appropriate expert processing units to features at different scales. Additionally, we design a Multi-scale Feature Refinement Module (MSFR) that captures multi-scale context through a combination of channel-spatial attention mechanisms and depthwise separable convolution, and gradually eliminates feature distribution differences between modalities using a dual-path residual learning strategy. Abundant experimental findings verify that the proposed CMFNet outperforms the 23 existing State-of-the-art (SOTA) methods.
Keywords:
Cross-fusion
mixture-of-experts
multi-scale
salient object detection
Journal
IF:
9.7
Papers:
4.4K
Citations:
2.4W
