Return
Cross-Modal Bayesian Inference for training-free open-vocabulary object detection in remote sensing images
Y
Y
Z
X
C
Q
DOI:10.1016/j.knosys.2026.116774.png)
Abstract
En 中文
Open-vocabulary object detection aims to localize and recognize objects from an open category space without being restricted to predefined classes. While recent training-free methods leverage powerful pretrained foundation models to avoid costly detector training, their performance in remote sensing scenarios remains limited due to category hallucination, insufficient category awareness, and unreliable candidate predictions. To address these challenges, CMBI is proposed as a Cross-Modal Bayesian Inference for training-free open-vocabulary object detection in remote sensing images as a rigorous Bayesian Maximum A Posteriori (MAP) estimation problem. Integrating SAM3 as the base detector and Qwen3 as the semantic assistant, CMBI performs collaborative reasoning through three probabilistically grounded components. First, a global semantic prior formulation mechanism is introduced to tightly constrain the open hypothesis space, effectively suppressing category hallucinations. Second, an adaptive conditioned likelihood estimation strategy employs dynamic visual anchors alongside text to sharpen spatial probability distributions, enhancing the perception of hard-to-recognize targets. Third, a posterior calibration and MAP decision module extracts local semantic evidence to iteratively calibrate intermediate probabilities, thereby eliminating spatial grouping ambiguities and facilitating more robust, bounding box predictions. Extensive experiments on three representative remote sensing benchmarks, DIOR, NWPU VHR-10, and HRRSD, demonstrate that CMBI consistently achieves superior performance over existing training-free baselines. These results validate the effectiveness of cross-modal collaboration for improving the accuracy and robustness of training-free open-vocabulary object detection in remote sensing imagery.
Keywords:
Object detection
Open-vocabulary
Training-free
Foundation models
Deep learning
Journal
K
IF:
7.6
Papers:
1.2W
Citations:
4.5W
