arrow
Return

Encoder-Decoder Architectures based Video Summarization using Key-Shot Selection Model

delete2023-09-16
delete4
PRE
AI
K
K. Yashwanth
B
Badal Soni *
DOI:10.1007/s11042-023-16700-3delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
With the exponential growth of video data, video summarization has become a challenging task. In this article, we propose a deep learning framework for video summarization that utilizes a sequence learning cum encoder-decoder network architecture with a key-shot selection model. We develop two RNN-based deep models, Additive Attentive Summariser (AAS) and Multiplicative Attentive Summariser (MAS), as well as a CNN-based model named - Sequential CNN Summariser (SCS). Our SCS and MAS model displays state-of-the-art performance in semantic segmentation, which we leverage to achieve superior performance in video summarization. We evaluate our models on two well-known datasets, SumMe and TVSum, and show that our proposed MAS and SCS models outperform state-of-the-art models such as DR-DSN. The proposed MAS model achieved an average F1 score of 44.1% and 60.7% on SumMe and TVSum datasets, respectively. Further, our contributions include the development of novel RNN-based and CNN-based models for video summarization and comprehensive experimental evaluations on multiple datasets that demonstrate the effectiveness of our proposed models.
Keywords:
Deep Learning
CNN
LSTM
RNN
Key-Shot
Summarization

Journal

Multimedia Tools and Applications cover
Multimedia Tools and Applications
IF:
3
Papers:
1.9W
Citations:
3.2W

Organization

N
national institute of technology (nit system)
Scholars:
4.0W
Papers: 3.7W
Citations: 31