arrow
Return

Cross-modal contrastive learning-based object detection under incomplete modalities

delete2026-04-07
delete0
delete
OA
AI
H
Hongjun Ma
倪欢 (Huan Ni) *
Y
Yongshi Jie
H
Haiyan Guan *
Q
Qingli Luo
DOI:10.1080/10095020.2026.2633014delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Object detection demonstrates stronger performance when using multimodal remote-sensing images than when using single-modal data. However, in practical applications, some modalities may be unavailable in a specific area, which limits the application of multimodal object detection methods. To address this challenge, a cross-modal contrastive learning and knowledge distillation (CCLKD) method is proposed in this paper. CCLKD is composed of dual branches for both easy-to-detect and hard-to-detect modalities. When CCLKD is training, it employs the knowledge distillation strategy to enhance the performance of the hard-to-detect branch (student network) by transferring the knowledge from the easy-to-detect branch (teacher network). As a result, when the easy-to-detect modality is absent, CCLKD can obtain good detection performance using only hard-to-detect modal data. To more effectively enhance the representation ability of hard-to-detect modality, CCLKD introduces an adaptive temperature-based knowledge distillation (ATKD) strategy and a category-constrained contrastive learning (CCL) mechanism. ATKD dynamically adjusts the distillation temperature based on the predicted probability provided by the teacher network and provides logic-, feature-, and relationship-level distillations. CCL strengthens the similarity between instances of the same category while suppressing interference from different categories, thereby improving intraclass compactness and interclass separability in the feature space. We employed three standard object detection datasets and compared CCLKD with state-of-the-art methods to validate its performance. The experimental results demonstrate the effectiveness of each component within CCLKD and further validate its superiority over existing methods through both quantitative and visual comparisons.
Keywords:
Feature alignment
object detection
contrastive learning
category-level information

Journal

G
Geo-Spatial Information Science
IF:
5.5
Papers:
838
Citations:
2.4K

Organization

T
tianjin university
Scholars:
8.0W
Papers: 5.7W
Citations: 88
N
Nanjing University of Information Science and Technology
Scholars:
2.8K
Papers: 1.2K
Citations: 1.7W
researcher View more organizations