Return
Modality Adaptive Representation Network for Efficient RGB-T Semantic Segmentation
DOI:10.1109/LSP.2025.3619880.png)
Abstract
En 中文
Due to the modality discrepancies caused by distinct imaging mechanisms, how to extract discriminative features from different modalities is always regarded as one of the key challenges in RGB-T semantic segmentation. Existing RGB-T semantic segmentation methods primarily employ feature extractors with dual-stream or single-stream structures to capture RGB and thermal features. This still struggles to effectively extract multimodal features with high discriminability while fully maintaining parameter efficiency. To address this issue, in this paper, a novel Modality Adaptive Representation Network (MARNet) is presented for efficient RGB-T semantic segmentation. Specifically, a plug-and-play Modality Adaptive Representation (MAR) strategy is designed to extract discriminative RGB and thermal features in a parameter-efficient manner by introducing learnable modality-specific prompts with self-modal and cross-modal image reconstruction constraints upon the feature extractor with single-stream structure. Moreover, a Multi-level Interaction Fusion (MIF) module is proposed to achieve multimodal feature interaction and fusion at the token, regional and global levels. Extensive experimental results on two public datasets demonstrate that, compared to mainstream methods, our proposed MARNet achieves competitive segmentation performance but with lower number of parameters.
Keywords:
Feature extraction
Image reconstruction
Semantic segmentation
Periodic structures
Training
Semantics
Radio frequency
Interpolation
Convolution
Adaptive systems
RGB-T semantic segmentation
multimodal feature extraction
modality-specific prompts
parameter efficiency

