Return
Action Localization Using 2D-CNN and 3D-CNN Collaboration
DOI:10.1109/ACCESS.2022.3193158.png)
Abstract
En 中文
Detecting human actions in videos is crucial for human-computer interaction, intelligent security, etc. Temporal and spatial information is very important for action localization, as how we utilize them determines the performance of detection. In order to address the problems of poor real-time and complex models of human action localization, a real-time detection architecture with two branches is presented in this paper, which predicts bounding boxes and action probabilities directly from video clips in one evaluation. We adopt the similar guidelines of YOLO (You Only Look Once) for bounding box regression and classification. The difference is that our model completes two tasks through two branches, one branch is used to extract spatial information from a key frame for bounding box regression, and another branch extracts spatiotemporal information from consecutive video frames for classification. In this paper, our model provides 65 frames-per-second on 16-frames input clips. Results on UCF-Sports and JHMDB-21 datasets present comparable accuracy to State-of-the-Art approaches. We obtain F-mAP of 90.39% and 76.29% with gains of 2.19% and 1.89%, respectively. On the JHMDB-21dataset, V-mAP reaches 90.7%, 90.0% and 69.5% with gains of 2.9%, 4.3% and 11.4% at IoU thresholds of 0.2, 0.5, and 0.75.
Keywords:
Feature extraction
Three-dimensional displays
Licenses
Data mining
Location awareness
Spatiotemporal phenomena
Shape
Action localization
real-time performance
deep learning
Journal
IF:
3.6
Papers:
9.8W
Citations:
29.4W
Organization
No organization information available

