arrow
Return

Action Keypoint Network for Efficient Video Recognition

delete2022-01-01
delete3
delete
OA
AI
X
Xu Chen
Y
Yahong Han *
X
Xiaohan Wang
Y
Yifan Sun
Y
Yi Yang
DOI:10.1109/TIP.2022.3191461delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Reducing redundancy is crucial for improving the efficiency of video recognition models. An effective approach is to select informative content from the holistic video, yielding a popular family of dynamic video recognition methods. However, existing dynamic methods focus on either temporal or spatial selection independently while neglecting a reality that the redundancies are usually spatial and temporal, simultaneously. Moreover, their selected content is usually cropped with fixed shapes (e.g., temporally-cropped frames, spatially-cropped patches), while the realistic distribution of informative content can be much more diverse. With these two insights, this paper proposes to integrate temporal and spatial selection into an Action Keypoint Network (AK-Net). From different frames and positions, AK-Net selects some informative points scattered in arbitrary-shaped regions as a set of action keypoints and then transforms the video recognition into point cloud classification. More concretely, AK-Net has two steps, i.e., the keypoint selection and the point cloud classification. First, it inputs the video into a baseline network and outputs a feature map from an intermediate layer. We view each pixel on this feature map as a spatial-temporal point and select some informative keypoints using self-attention. Second, AK-Net devises a ranking criterion to arrange the keypoints into an ordered 1D sequence. Since the video is represented with a 1D sequence after the specified layer, AK-Net transforms the subsequent layers into a point cloud classification sub-net by compacting the original 2D convolutional kernels into 1D kernels. Consequentially, AK-Net brings two-fold benefits for efficiency: The keypoint selection step collects informative content within arbitrary shapes and increases the efficiency for modeling spatial-temporal dependencies, while the point cloud classification step further reduces the computational cost by compacting the convolutional kernels. Experimental results show that AK-Net can consistently improve the efficiency and performance of baseline methods on several video recognition benchmarks.
Keywords:
Video recognition
space-time interest points
deep learning
point cloud

Journal

IEEE Transactions on Image Processing cover
IEEE Transactions on Image Processing
IF:
13.7
Papers:
1.0W
Citations:
8.4W

Organization

T
tianjin university
Scholars:
8.0W
Papers: 5.8W
Citations: 88
B
baidu
Scholars:
578
Papers: 471
Citations: 1
Z
zhejiang university
Scholars:
17.7W
Papers: 12.1W
Citations: 152
researcher View more organizations
Cited Papers

Cited Papers

err
IF0
err
err0
PREAI
err
errShare
errSave
errShare
errSave
Introducing landscape ecology
err1987-07-01
err0
PREAI
errFrank B. Golley
errShare
errSave
Time-dependent depolarization of aligned HD molecules
err2009-01-01
err0
errOAAI
errNate C.-M. Bartlett; Daniel J. Miller; Richard N. Zare; Andrew J. Alexander; Dimitris Sofikitis; T. Peter Rakitzis
errShare
errSave
Sequential Video VLAD: Training the Aggregation Locally and Temporally
err2018-10-01
err91
PREAI
errXu, Youjiang; Han, Yahong; Hong, Richang; Tian, Qi
errShare
errSave
The mental health of staff working in intensive care during COVID-19
err
IF0
err2020-11-04
err0
errOAAI
errNeil Greenberg; Dale Weston; Charlotte Hall; Tristan Caulfield; Victoria Williamson; Kevin Fong
errShare
errSave
researcher View more