Return
Multimodal collaborative saliency object detection network using MCSDNet
DOI:10.1016/j.displa.2026.103467.png)
Abstract
En 中文
• Novel three-stream architecture optimizing CNN, CapsNet, and boundary-guided network for feature synergy. • Dynamic cross-modal interaction (SIM/CAM) enabling semantic alignment and noise suppression. • Achieved top performance in 13 out of 20 saliency metrics, and ranked first on four datasets for boundary-related metrics, with improvements of 3%, 2.7%, 0.7%, and 2.9% over the second-best method. • Runs at 22.4 FPS on an NVIDIA GeForce RTX 4080 SUPER (16 GB) GPU, ranking third among all compared methods.
Keywords:
Multimodal collaboration
Saliency detection
CNN
CapsNet
Boundary-guided network
Journal
IF:
3.4
Papers:
2.1K
Citations:
3.2K

