arrow
返回

Deep multi-scale pyramidal features network for supervised video summarization

delete2024-03-01
delete29
delete
OA
AI
H
Habib Khan
T
Tanveer Hussain
S
Samee U. Khan
Z
Zulfiqar Ahmad Khan
S
Sung Wook Baik *
DOI:10.1016/j.eswa.2023.121288delete
delete原文链接
delete分享
delete收藏
查看原文
摘要

摘要

En 中文
Video data are witnessing exponential growth, and extracting summarized information is challenging. It is always necessary to reduce the load of video traffic for the efficient video storage, transmission, and retrieval requirements. The aim of video summarization (VS) is to extract the most important contents from video repositories effectively. Recent attempts have used fewer representative features, which have been fed to recurrent networks to achieve VS. However, generating the desired summaries can become challenging due to the limited representativeness of extracted features and a lack of consideration for feature refinement. In this article, we introduce a vision transformer (ViT)-assisted deep pyramidal refinement network that can extract and refine multi-scale features and can predict an importance score for each frame. The proposed network comprises four main modules; initially, a dense prediction transformer with a ViT backbone is applied for the first time in this domain to acquire the optimal representations from the input frames. Then, feature maps are obtained from various layers separately and processed individually to support multi-scale progressive feature fusion and refinement before the data are passed to the ultimate prediction module. Next, a customized pyramidal refinement block is employed to refine the multi-level feature set before predicting the importance scores. Finally, video summaries are produced by selecting keyframes based on the predictions. To explore the performance of the proposed network, extensive experiments are conducted on the TVSum and SumMe datasets, and our network is found to achieve F1-scores of 62.4% and 51.9%, respectively, outperforming state-of-the-art alternatives by 0.9% and 0.5%.
Keyword:
Video summarization
Supervised learning
Keyframes
Feature fusion
Keyshots
Feature refinement
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

Expert Systems with Applications 封面图
Expert Systems with Applications
IF:
7.5
论文数:
2.9W
被引数:
10.2W

机构

S
Sejong University
学者数:
8.3K
论文数: 1.1W
被引数: 1.5W
E
Edge Hill University
学者数:
1.2K
论文数: 1.3K
被引数: 958