arrow
返回

Multi-scale deep feature fusion based sparse dictionary selection for video summarization

delete2023-10-01
delete2
PRE
AI
X
Xiao Wu
M
Mingyang Ma
S
Shuai Wan
S
Shaohui Mei *
DOI:10.1016/j.image.2023.117006delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
The explosive growth of video data constitutes a series of new challenges in computer vision, and the function of video summarization (VS) is becoming more and more prominent. Recent works have shown the effectiveness of sparse dictionary selection (SDS) based VS, which selects a representative frame set to sufficiently reconstruct a given video. Existing SDS based VS methods use conventional handcrafted features or single-scale deep features, which could diminish their summarization performance due to the underutilization of frame feature representation. Deep learning techniques based on convolutional neural networks (CNNs) exhibit powerful capabilities among various vision tasks, as the CNN provides excellent feature representation. Therefore, in this paper, a multi-scale deep feature fusion based sparse dictionary selection (MSDFF-SDS) is proposed for VS. Specifically, multi-scale features include the directly extracted features from the last fully connected layer and the global average pooling (GAP) processed features from intermediate layers, then VS is formulated as a problem of minimizing the reconstruction error using the multi-scale deep feature fusion. In our formulation, the contribution of each scale of features can be adjusted by a balance parameter, and the row-sparsity consistency of the simultaneous reconstruction coefficient is used to select as few keyframes as possible. The resulting MSDFF-SDS model is solved by using an efficient greedy pursuit algorithm. Experimental results on two benchmark datasets demonstrate that the proposed MSDFF-SDS improves the F-score of keyframe based summarization more than 3% compared with the existing SDS methods, and performs better than most deep-learning methods for skimming based summarization.
Keyword:
Video summarization
Keyframe
Sparse coding
Dictionary selection
Multi-scale

期刊

S
Signal Processing and Image Communication
IF:
2.7
论文数:
2.8K
被引数:
4.2K

机构

X
xi'an jiaotong university
学者数:
9.3W
论文数: 6.7W
被引数: 75
N
Northwestern Polytechnical University
学者数:
4.6W
论文数: 3.7W
被引数: 5.3W
引用论文

引用论文

Demographic estimates of hunter–gatherers during the Last Glacial Maximum in Europe against the background of palaeoenvironmental data
err2016-12-01
err0
errOAAI
errAndreas Maier; Frank Lehmkuhl; Patrick Ludwig; Martin Melles; Isabell Schmidt; Yaping Shao; Christian Zeeden; Andreas Zimmermann
err分享
err收藏
Effects of Novel Isoform-Selective Phosphoinositide 3-Kinase Inhibitors on Natural Killer Cell Function
err2014-06-10
err0
errOAAI
errSung Su Yea; Lomon So; Sharmila Mallya; Jongdae Lee; Kamalakannan Rajasekaran; Subramaniam Malarkannan; David A. Fruman
err分享
err收藏
Adaptive Greedy Dictionary Selection for Web Media Summarization
err2017-01-01
err50
PREAI
errCong, Yang; Liu, Ji; Sun, Gan; You, Quanzeng; Li, Yuncheng; Luo, Jiebo
err分享
err收藏
学者 查看更多内容