1
Return

DSFusion: different size modalities zero-shot segmentation via heterogeneous fusion

delete2026-07-30
delete0
PRE
AI
Y
Yue Zhuo
D
Di Zhou *
P
Pengpeng Xu
S
Shilun Liu
Y
Yan Tian *
DOI:10.1007/s00530-026-02523-3delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Zero-shot segmentation based on RGB-D data plays a crucial role in embodied intelligence systems and autonomous driving technologies. However, current approaches face challenges with heterogeneous data fusion in multimodal foundation model (MFM)-based methods because the original segment anything model (SAM) is designed only for images or videos. In addition, the research on the fusion of multimedia data of different sizes is omitted. Motivated by the memory mechanism in zero-shot segmentation, we design DSFusion, an approach for heterogeneous data fusion in zero-shot segmentation, where multimodal data are considered as a modality sequence to capture the modality-agnostic feature in the revised memory mechanism. In addition, the large-size modality is divided and assembled into channels, avoiding both loss of detail and noise introduced during upsampling. The experimental results in the ScanNet V2 and ScanNet200 datasets indicate that our approach improves the mean intersection over union (mIoU) by a margin of 3.33% and 3.42% when compared to the prevailing approaches. Project page: https://faith643.github.io/SAM2-Based_RGB-D_Zero-Shot_Segmentation
Keywords:
RGB-D data
Heterogeneous data fusion
Semantic segmentation
Segment anything model

Journal

Multimedia Systems cover
Multimedia Systems
IF:
3.1
Papers:
2.7K
Citations:
2.7K

Organization

S
School of Micro-Nano Electronics
Scholars:
2
Papers: 2
Citations: 0
S
School of Computer Science and Technology
Scholars:
1.3K
Papers: 514
Citations: 0
S
school of cyberspace
Scholars:
12
Papers: 4
Citations: 0
Cited Papers

Cited Papers

Citing Papers

Citing Papers