Return
CM-Mamba: State–space cross-modal fusion for CT and chest X-ray lung disease classification
G
S
DOI:10.1016/j.bspc.2026.111211.png)
Abstract
En 中文
Accurate multimodal fusion of chest computed tomography (CT) and chest X-ray (CXR) images remains challenging because of the heterogeneous characteristics of the two imaging modalities and the limited availability of patient-aligned multimodal datasets. This paper presents CM-Mamba, a lightweight state–space-based cross-modal fusion framework for lung disease classification. The proposed framework employs modality-specific linear-time Mamba encoders and introduces a novel two-stage fusion strategy that performs Mamba-Enhanced Cross-Attention Fusion (MECAF) before Modality-Adaptive Feature Harmonization (MAFH), enabling complementary cross-modal feature interaction prior to adaptive feature refinement. CM-Mamba is developed using two publicly available unpaired, class-matched CT–CXR datasets and further validated on an independent paired clinical CT–CXR dataset comprising COVID-19, pneumonia, and normal cases. Under five-fold cross-validation, the proposed framework achieves mean classification accuracies of 99.48% and 99.50% on the two public datasets, respectively, while demonstrating strong generalization on the paired clinical dataset. Ablation results show that applying MECAF before MAFH improves accuracy by 0.83%. CM-Mamba also improves accuracy by 4.44% and 1.53% over the CXR-only and CT-only models, respectively. The complete framework requires 17.16M parameters and 5.38 GFLOPs, while the lightweight MECAF and MAFH modules add minimal overhead, enabling an effective balance between diagnostic performance and computational efficiency. The implementation of CM-Mamba is publicly available at https://github.com/gautamiphd-design/CM-Mamba .
Journal
IF:
4.9
Papers:
9.7K
Citations:
2.4W
Organization
U
