Return
CODNet: Context-based object detection network for multimodal image captioning and virtual question answering
DOI:10.1016/j.imavis.2025.105768.png)
Abstract
En 中文
• Context-based object detection refers to identifying objects within an image or a sequence of images by considering the surrounding context in which these objects appear. For example, a keyboard is usually found near a computer. Some objects are more likely to appear together. For example, a fork and a knife are often found together on a dining table. • This study presents a new context-based object detection network that defines a new direction in visual object detection and improves LLMs. • A novel CODNet network is proposed that helps in first generating the object words and then classifying the objects in the images. The network comprises three sub-modules: encoder, Llama, and decoder. The network combines CNN-mLSTM (Convolutional Neural Network merged with multiplicative LSTM) to help extract features from provided images. Llama is a pre-trained large language model for decoding the context and finetuned YOLOv6 for object detection. • A new set of object words, namely, ERCODE (Easy Readable Context-Based Object DEtection), is also proposed in this work to help classify required objects in images.
Keywords:
CNN
Large language model
Multiplicative LSTM
Multimodal modeling
Yolov6
AI Summary
Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.
Journal
IF:
4.2
Papers:
4.0K
Citations:
6.7K

