arrow
Return

Squacc BiLSTM: a framework for dense video captioning using neural knowledge graph and deep learning

delete2025-09-09
delete1
PRE
AI
H
Hugo Terashima‐Marín
P
Peyman Najafirad
S
Santiago Enrique Conant-Pablos
M
Mohd Anas Wajid
DOI:10.1007/s11760-025-04657-9delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En
Dense Video Captioning narrates a thorough comprehension of the video through accurate identification of the frame details, including the objects and their movements. However, most of the existing models fail to learn the visual context information and linguistic clues present in the video, which may limit the model's cognitive capacity to produce the descriptions. Therefore, in this research, the automatic dense video captioning is performed using the Squacc Bidirectional Long Short-Term Memory (Squacc BiLSTM) model, where a Neural knowledge graph (NKG) is generated based on the recurrent neural network. The generated NKG unveils the temporal video features, which support. The prediction of the video contents for captioning using Squacc BiLSTM classifier. Furthermore, the Squacc optimization algorithm fine-tunes the classifier parameters, which supports capturing the past and future contexts of the video for precise captioning. The experimental results demonstrate that the proposed Squacc BiLSTM model has been proven effective in video captioning, showcasing enhanced BLEU, ROUGE, CIDEr, METEOR, and SPICE scores of 0.439, 0.511, 0.759, 0.264, and 19.994, outperforming the existing techniques.
Keywords:
Dense video captioning
Deep learning
Neural knowledge graph
Optimal solution
Reliable

Journal

Signal Image and Video Processing cover
Signal Image and Video Processing
IF:
2.1
Papers:
877
Citations:
4.6K

Organization

No organization information available