arrow
Return

Competitive fusion in multimodal networks for enhanced salient object detection

delete2026-07-09
delete0
PRE
AI
H
Hanzhong Tan
S
Shuangbing Wen
L
Lingfeng Zhang
J
Jun Li
T
Tao Hu *
DOI:10.1007/s00371-026-04602-ydelete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Salient object detection (SOD) aims to identify the most visually compelling regions in images, playing a crucial role in computer vision tasks. Traditional RGB-D/T multimodal interaction methods such as Hypersor with multiplicative operations, FasterSal via concatenation, and SACNet using equal interaction attention typically rely on equal interaction mechanisms and convolutional computations, exhibiting limitations in dynamically and adaptively capturing and representing critical discriminative features. For salient object detection, this paper introduces MC2FNet, a multi-scale cross-modal competitive fusion network for RGB and depth/thermal data, designed to achieve efficient intra-modal, cross-modal, and multi-scale competitive fusion. MC2FNet employs a hybrid attentional approach to reduce computational overhead and enhance feature extraction. Experimental results on RGB-T and RGB-D datasets demonstrate that MC2FNet achieves state-of-the-art performance, outperforming existing methods in accuracy and efficiency. This work advances the field by providing a novel framework for robust multimodal salient object detection. The code will be available at https://github.com/liangjiaxiaoqi/MC2FNet .
Keywords:
Salient object detection
multimodal competitive fusion
Group aggregation and re-embedding
Spatial synergistic suppression–enhancement

Journal

T
The Visual Computer
IF:
0
Papers:
369
Citations:
0

Organization

C
college of mathematics and statistic
Scholars:
2
Papers: 1
Citations: 0
researcher View more organizations