1
Return

Cross-Modal Bayesian Inference for training-free open-vocabulary object detection in remote sensing images

delete2026-08-03
delete0
PRE
AI
Y
Yan Li
Y
Yunpeng Bai *
Z
Zhenhua Wu
X
Xingguo Zhang
李颖 cover
李颖 (Ying Li)
C
Changjing Shang
Q
Qiang Shen
DOI:10.1016/j.knosys.2026.116774delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Open-vocabulary object detection aims to localize and recognize objects from an open category space without being restricted to predefined classes. While recent training-free methods leverage powerful pretrained foundation models to avoid costly detector training, their performance in remote sensing scenarios remains limited due to category hallucination, insufficient category awareness, and unreliable candidate predictions. To address these challenges, CMBI is proposed as a Cross-Modal Bayesian Inference for training-free open-vocabulary object detection in remote sensing images as a rigorous Bayesian Maximum A Posteriori (MAP) estimation problem. Integrating SAM3 as the base detector and Qwen3 as the semantic assistant, CMBI performs collaborative reasoning through three probabilistically grounded components. First, a global semantic prior formulation mechanism is introduced to tightly constrain the open hypothesis space, effectively suppressing category hallucinations. Second, an adaptive conditioned likelihood estimation strategy employs dynamic visual anchors alongside text to sharpen spatial probability distributions, enhancing the perception of hard-to-recognize targets. Third, a posterior calibration and MAP decision module extracts local semantic evidence to iteratively calibrate intermediate probabilities, thereby eliminating spatial grouping ambiguities and facilitating more robust, bounding box predictions. Extensive experiments on three representative remote sensing benchmarks, DIOR, NWPU VHR-10, and HRRSD, demonstrate that CMBI consistently achieves superior performance over existing training-free baselines. These results validate the effectiveness of cross-modal collaboration for improving the accuracy and robustness of training-free open-vocabulary object detection in remote sensing imagery.
Keywords:
Object detection
Open-vocabulary
Training-free
Foundation models
Deep learning

Journal

K
Knowledge-Based Systems
IF:
7.6
Papers:
1.2W
Citations:
4.5W

Organization

T
test technology institute
Scholars:
2
Papers: 1
Citations: 0
N
northwestern polytechnical university
Scholars:
1.0W
Papers: 3.8K
Citations: 0
A
Aberystwyth University
Scholars:
2.5K
Papers: 2.5K
Citations: 4.3K
Cited Papers

Cited Papers

Citing Papers

Citing Papers