arrow
返回

Action Keypoint Network for Efficient Video Recognition

delete2022-01-01
delete3
delete
OA
AI
X
Xu Chen
Y
Yahong Han *
X
Xiaohan Wang
Y
Yifan Sun
Y
Yi Yang
DOI:10.1109/TIP.2022.3191461delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Reducing redundancy is crucial for improving the efficiency of video recognition models. An effective approach is to select informative content from the holistic video, yielding a popular family of dynamic video recognition methods. However, existing dynamic methods focus on either temporal or spatial selection independently while neglecting a reality that the redundancies are usually spatial and temporal, simultaneously. Moreover, their selected content is usually cropped with fixed shapes (e.g., temporally-cropped frames, spatially-cropped patches), while the realistic distribution of informative content can be much more diverse. With these two insights, this paper proposes to integrate temporal and spatial selection into an Action Keypoint Network (AK-Net). From different frames and positions, AK-Net selects some informative points scattered in arbitrary-shaped regions as a set of action keypoints and then transforms the video recognition into point cloud classification. More concretely, AK-Net has two steps, i.e., the keypoint selection and the point cloud classification. First, it inputs the video into a baseline network and outputs a feature map from an intermediate layer. We view each pixel on this feature map as a spatial-temporal point and select some informative keypoints using self-attention. Second, AK-Net devises a ranking criterion to arrange the keypoints into an ordered 1D sequence. Since the video is represented with a 1D sequence after the specified layer, AK-Net transforms the subsequent layers into a point cloud classification sub-net by compacting the original 2D convolutional kernels into 1D kernels. Consequentially, AK-Net brings two-fold benefits for efficiency: The keypoint selection step collects informative content within arbitrary shapes and increases the efficiency for modeling spatial-temporal dependencies, while the point cloud classification step further reduces the computational cost by compacting the convolutional kernels. Experimental results show that AK-Net can consistently improve the efficiency and performance of baseline methods on several video recognition benchmarks.
Keyword:
Video recognition
space-time interest points
deep learning
point cloud

期刊

IEEE Transactions on Image Processing 封面图
IEEE Transactions on Image Processing
IF:
13.7
论文数:
1.0W
被引数:
8.4W

机构

T
tianjin university
学者数:
8.0W
论文数: 5.8W
被引数: 88
B
baidu
学者数:
578
论文数: 471
被引数: 1
Z
zhejiang university
学者数:
17.7W
论文数: 12.1W
被引数: 152
学者 查看更多机构
引用论文

引用论文

err
IF0
err
err0
PREAI
err
err分享
err收藏
err分享
err收藏
Introducing landscape ecology
err1987-07-01
err0
PREAI
errFrank B. Golley
err分享
err收藏
Time-dependent depolarization of aligned HD molecules
err2009-01-01
err0
errOAAI
errNate C.-M. Bartlett; Daniel J. Miller; Richard N. Zare; Andrew J. Alexander; Dimitris Sofikitis; T. Peter Rakitzis
err分享
err收藏
Sequential Video VLAD: Training the Aggregation Locally and Temporally
err2018-10-01
err91
PREAI
errXu, Youjiang; Han, Yahong; Hong, Richang; Tian, Qi
err分享
err收藏
The mental health of staff working in intensive care during COVID-19
err
IF0
err2020-11-04
err0
errOAAI
errNeil Greenberg; Dale Weston; Charlotte Hall; Tristan Caulfield; Victoria Williamson; Kevin Fong
err分享
err收藏
学者 查看更多内容