arrow
返回

Deep Unsupervised Key Frame Extraction for Efficient Video Classification

delete2023-02-25
delete17
delete
OA
AI
H
Hao Tang *
L
Lei Ding
S
Songsong Wu
B
Bin Ren
N
Nicu Sebe
P
Paolo Rota
DOI:10.1145/3571735delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Video processing and analysis have become an urgent task, as a huge amount of videos (e.g., YouTube, Hulu) are uploaded online every day. The extraction of representative key frames from videos is important in video processing and analysis since it greatly reduces computing resources and time. Although great progress has been made recently, large-scale video classification remains an open problem, as the existing methods have not well balanced the performance and efficiency simultaneously. To tackle this problem, this work presents an unsupervised method to retrieve the key frames, which combines the convolutional neural network and temporal segment density peaks clustering. The proposed temporal segment density peaks clustering is a generic and powerful framework, and it has two advantages compared with previous works. One is that it can calculate the number of key frames automatically. The other is that it can preserve the temporal information of the video. Thus, it improves the efficiency of video classification. Furthermore, a long short-term memory network is added on the top of the convolutional neural network to further elevate the performance of classification. Moreover, a weight fusion strategy of different input networks is presented to boost performance. By optimizing both video classification and key frame extraction simultaneously, we achieve better classification performance and higher efficiency. We evaluate our method on two popular datasets (i.e., HMDB51 and UCF101), and the experimental results consistently demonstrate that our strategy achieves competitive performance and efficiency compared with the state-of-the-art approaches.
Keyword:
Key frame extraction
density peaks clustering
LSTM
weight fusion
unsupervised learning
video classification

期刊

ACM Transactions on Multimedia Computing Communications and Applications 封面图
ACM Transactions on Multimedia Computing Communications and Applications
IF:
6
论文数:
2.0K
被引数:
5.4K

机构

U
University of Trento
学者数:
8.8K
论文数: 9.0K
被引数: 1.2W
E
ETH Zurich
学者数:
3.0W
论文数: 2.4W
被引数: 8.4W
S
swiss federal institutes of technology domain
学者数:
9.0W
论文数: 8.0W
被引数: 163
学者 查看更多机构
引用论文

引用论文

Climate policy: Steps to China's carbon peak气候政策: 迈向中国碳峰值的步骤
err2015-06-17
err0
errOAAI
errZhu Liu; Dabo Guan; Scott Moore; Henry Lee; Jun Su; Qiang Zhang
err分享
err收藏
err分享
err收藏
High-resolution 13C n.m.r. spectra of solid nitrogen-containing compounds
err1980-01-01
err0
PREAI
errChristopher J. Groombridge; Robin K. Harris; Kenneth J. Packer; Barry J. Say; Steven F. Tanner
err分享
err收藏
ImageNet Large Scale Visual Recognition ChallengeImageNet大规模视觉识别挑战
err2015-04-11
err2.7W
PREAI
errRussakovsky, Olga; Deng, Jia; Su, Hao; Krause, Jonathan; Satheesh, Sanjeev; Ma, Sean; Huang, Zhiheng; Karpathy, Andrej; Khosla, Aditya; Bernstein, Michael; Berg, Alexander C.; Fei-Fei, Li
err分享
err收藏
Representative Selection with Structured Sparsity
err2017-03-01
err46
errOAAI
errWang, Hongxing; Kawahara, Yoshinobu; Weng, Chaoqun; Yuan, Junsong
err分享
err收藏
Gas Turbine Performance燃气轮机性能
err
IF0
err2008-02-11
err0
PREAI
errPhilip P. Walsh; Paul Fletcher
err分享
err收藏
Positioning in cellular networks: Past, present, future
err2018-04-01
err0
PREAI
errSara Modarres Razavi; Fredrik Gunnarsson; Henrik Ryden; Ake Busin; Xingqin Lin; Xin Zhang; Satyam Dwivedi; Iana Siomina; Ritesh Shreevastav
err分享
err收藏
学者 查看更多内容