arrow
Return

Bidirectional transformer with knowledge graph for video captioning

delete2023-12-21
delete1
PRE
AI
钟茂生 (Maosheng Zhong)
Y
Youde Chen *
H
Hao Zhang
H
Hao Xiong
W
Wang, Zhixiang
DOI:10.1007/s11042-023-17822-4delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Models based on transformer architecture have risen to prominence for video captioning. However, most models are only to improve either the encoder or the decoder, because when we improve the encoder and decoder simultaneously, the shortcomings of either side may be amplified. Based on the transformer architecture, we connect a bidirectional decoder and an encoder that integrates fine-grained spatio-temporal features, objects, and relationships between the objects in the video. Experiments show that improvements in the encoder amplify the information leakage of the bidirectional decoder and further produce a worse result. To tackle this problem, we generate pseudo reverse captions and propose a Bidirectional Transformer with Knowledge Graph (BTKG), which integrates the outputs of two encoders into the forward and backward decoders of the bidirectional decoder, respectively. In addition, we make fine-grained improvements on the interior of the different encoders according to four modal features of the video. Experiments on two mainstream benchmark datasets, i.e., MSVD and MSR-VTT, demonstrate the effectiveness of BTKG, which achieves state-of-the-art performance in significant metrics. Moreover, the sentences generated by BTKG contain scene words and modifiers, that are more in line with human language habits. Codes are available on https://github.com/nickchen121/BTKG.
Keywords:
Video captioning
Bidirectional transformer
Knowledge graph
Multimodal of video

Journal

Multimedia Tools and Applications cover
Multimedia Tools and Applications
IF:
3
Papers:
1.9W
Citations:
3.2W

Organization

J
Jiangxi Normal University
Scholars:
6.9K
Papers: 4.7K
Citations: 8.8K