Return
Deep ensemble learning method for zero-shot object detection
DOI:10.1016/j.asoc.2025.114333.png)
Abstract
En 中文
• For the zero-shot learning (ZSD) model, we propose using the visual embedding space. This is because embedding features from the visual space into the semantic space can improve recognition. • With word vector representations, we propose to use the Random Multimodal Deep Learning Model (RMDL) to model textual descriptions using deep learning. • While our neural network model is highly adaptable and capable of learning the semantic space representation end-to-end, the focus of this paper is on establishing a highly effective multi-modality integration method. • The efficacy of our proposed method on benchmark datasets such as MSCOCO, Visual Genome, and ILSVRC 2012, among others, has been demonstrated through extensive experiments.
Journal
IF:
6.6
Papers:
1.4W
Citations:
4.8W

