arrow
返回

V2T: video to text framework using a novel automatic shot boundary detection algorithm

delete2022-03-08
delete8
PRE
AI
A
Alok Singh *
T
Thoudam Doren Singh
S
Sivaji Bandyopadhyay
DOI:10.1007/s11042-022-12343-ydelete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
The generation of natural language descriptions for a video has been reported by many researchers till now. But, it is still the most interesting research topic among the researchers due to the emerging interdisciplinary problem of Computer Vision (CV), Natural Language Processing (NLP) and Deep Learning (DL). The results of a video description are still not convincing due to the redundancy of a large number of similar frames in a video. In this paper, we propose dual-stage based text generation approach in which the first stage is for reducing redundancy due to the similar frames by processing selected sets of frames and keyframe from the shots of a video and in the second stage, the text generator module will generate relevant text for a video using the selected sets of frames and keyframes of each shot. In the first stage, a flexible novel shot boundary detection (SBD or temporal boundaries) approach is proposed which will segment the video into shots and then keyframe and set of frames are selected from each shot using frame selection policy. Then, the spatio-temporal features for each segment and 2D features for each keyframe are extracted respectively using the 3D convolutional network and VGG19. These features are passed to the next stage where these features are embedded with semantic concepts related to video and then text generation will take place using Long Short Term Memory (LSTM) recurrent network. The proposed approach is the amalgamation of classical and modern computer vision techniques. In the first stage, the Noise-Resistant Local Binary Pattern (NRLBP) feature is used for detecting illumination and motion invariant temporal boundaries in a video and processing keyframes and sets of frames for the further text generation. TRECVid 2001 and 2007 datasets are used to validate the exactness of the proposed SBD approach and MSR-VTT (Microsoft Research Video to Text ) and YouTube2text (MSVD) datasets are applied to analyze and validate the performance of proposed video to text generation approach.
Keyword:
Shot boundary detection
Illumination
Motion effect
Abrupt transition
Video captioning

期刊

Multimedia Tools and Applications 封面图
Multimedia Tools and Applications
IF:
3
论文数:
2.0W
被引数:
3.2W

机构

N
national institute of technology (nit system)
学者数:
4.0W
论文数: 3.7W
被引数: 31
引用论文

引用论文

Evaluation of the Mixing Effectiveness of a New Powder Mixer
err2008-10-20
err0
PREAI
errGiovanni Filippo Palmieri; Debora Lovato Leo Marchitto; Aldo Zanchetta; Sante Martelli
err分享
err收藏
Climate policy: Steps to China's carbon peak气候政策: 迈向中国碳峰值的步骤
err2015-06-17
err0
errOAAI
errZhu Liu; Dabo Guan; Scott Moore; Henry Lee; Jun Su; Qiang Zhang
err分享
err收藏
A Shot boundary Detection Technique based on Visual Colour Information
err2020-09-25
err18
PREAI
errChakraborty, Saptarshi; Thounaojam, Dalton Meitei; Sinha, Nidul
err分享
err收藏
Learning deep spatiotemporal features for video captioning
err2018-12-01
err11
PREAI
errDaskalakis, Eleftherios; Tzelepi, Maria; Tefas, Anastasios
err分享
err收藏
学者 查看更多内容