arrow
Return

CODNet: Context-based object detection network for multimodal image captioning and virtual question answering

delete2025-10-09
delete0
delete
OA
AI
C
Chhaya Gupta
N
Nasib Singh Gill
P
Preeti Gulia *
G
Giovanni Pau *
DOI:10.1016/j.imavis.2025.105768delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
• Context-based object detection refers to identifying objects within an image or a sequence of images by considering the surrounding context in which these objects appear. For example, a keyboard is usually found near a computer. Some objects are more likely to appear together. For example, a fork and a knife are often found together on a dining table. • This study presents a new context-based object detection network that defines a new direction in visual object detection and improves LLMs. • A novel CODNet network is proposed that helps in first generating the object words and then classifying the objects in the images. The network comprises three sub-modules: encoder, Llama, and decoder. The network combines CNN-mLSTM (Convolutional Neural Network merged with multiplicative LSTM) to help extract features from provided images. Llama is a pre-trained large language model for decoding the context and finetuned YOLOv6 for object detection. • A new set of object words, namely, ERCODE (Easy Readable Context-Based Object DEtection), is also proposed in this work to help classify required objects in images.
Keywords:
CNN
Large language model
Multiplicative LSTM
Multimodal modeling
Yolov6
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

Image and Vision Computing cover
Image and Vision Computing
IF:
4.2
Papers:
4.0K
Citations:
6.7K

Organization

M
Maharshi Dayanand University
Scholars:
1.9K
Papers: 1.6K
Citations: 2.1K
K
Kore University of Enna
Scholars:
69
Papers: 59
Citations: 2