arrow
Return

Multimodal attention-based transformer for video captioning

delete2023-07-09
delete4
PRE
AI
M
M. Hemalatha *
C
Charu Chandra
DOI:10.1007/s10489-023-04597-2delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Video captioning is a computer vision task that generates a natural language description for a video. In this paper, we propose a multimodal attention-based transformer using the keyframe features, object features, and semantic keyword embedding features of a video. The Structural Similarity Index Measure (SSIM) is used to extract keyframes from a video. We also detect the unique objects from the extracted keyframes. The features from the keyframes and the objects detected in the keyframes are extracted using a pretrained Convolutional Neural Network (CNN). In the encoder, we use a bimodal attention block to apply two-way cross-attention between the keyframe features and the object features. In the decoder, we combine the features of the words generated up to the previous time step, the semantic keyword embedding features, and the encoder features using a tri-modal attention block. This allows the decoder to choose the multimodal features dynamically to generate the next word in the description. We evaluated the proposed approach using the MSVD, MSR-VTT, and Charades datasets and observed that the proposed model provides better performance than other state-of-the-art models.
Keywords:
Video captioning
Transformer
Multimodal attention
Semantic keywords

Journal

Applied Intelligence cover
Applied Intelligence
IF:
3.5
Papers:
7.5K
Citations:
1.7W

Organization

I
indian institute of technology system (iit system)
Scholars:
9.5W
Papers: 9.9W
Citations: 93