arrow
返回

Multimodal attention-based transformer for video captioning

delete2023-07-09
delete4
PRE
AI
M
M. Hemalatha *
C
Charu Chandra
DOI:10.1007/s10489-023-04597-2delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Video captioning is a computer vision task that generates a natural language description for a video. In this paper, we propose a multimodal attention-based transformer using the keyframe features, object features, and semantic keyword embedding features of a video. The Structural Similarity Index Measure (SSIM) is used to extract keyframes from a video. We also detect the unique objects from the extracted keyframes. The features from the keyframes and the objects detected in the keyframes are extracted using a pretrained Convolutional Neural Network (CNN). In the encoder, we use a bimodal attention block to apply two-way cross-attention between the keyframe features and the object features. In the decoder, we combine the features of the words generated up to the previous time step, the semantic keyword embedding features, and the encoder features using a tri-modal attention block. This allows the decoder to choose the multimodal features dynamically to generate the next word in the description. We evaluated the proposed approach using the MSVD, MSR-VTT, and Charades datasets and observed that the proposed model provides better performance than other state-of-the-art models.
Keyword:
Video captioning
Transformer
Multimodal attention
Semantic keywords

期刊

Applied Intelligence 封面图
Applied Intelligence
IF:
3.5
论文数:
7.6K
被引数:
1.7W

机构

I
indian institute of technology system (iit system)
学者数:
9.5W
论文数: 9.9W
被引数: 93
引用论文

引用论文

err分享
err收藏
Data-Driven Deepfake Forensics Model Based on Large-Scale Frequency and Noise Features
err2024-01-01
err10
PREAI
errLan, Guipeng; Xiao, Shuai; Wen, Jiabao; Chen, Desheng; Zhu, Yong
err分享
err收藏
Climate policy: Steps to China's carbon peak气候政策: 迈向中国碳峰值的步骤
err2015-06-17
err0
errOAAI
errZhu Liu; Dabo Guan; Scott Moore; Henry Lee; Jun Su; Qiang Zhang
err分享
err收藏
学者 查看更多内容