arrow
Return

Action Localization Using 2D-CNN and 3D-CNN Collaboration

delete2022-01-01
delete2
delete
OA
AI
J
Jiale Tong
李建军 cover
李建军 (Jianjun Li) *
张明 cover
张明 (Ming Zhang)
张宝华 cover
张宝华 (Baohua Zhang)
DOI:10.1109/ACCESS.2022.3193158delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Detecting human actions in videos is crucial for human-computer interaction, intelligent security, etc. Temporal and spatial information is very important for action localization, as how we utilize them determines the performance of detection. In order to address the problems of poor real-time and complex models of human action localization, a real-time detection architecture with two branches is presented in this paper, which predicts bounding boxes and action probabilities directly from video clips in one evaluation. We adopt the similar guidelines of YOLO (You Only Look Once) for bounding box regression and classification. The difference is that our model completes two tasks through two branches, one branch is used to extract spatial information from a key frame for bounding box regression, and another branch extracts spatiotemporal information from consecutive video frames for classification. In this paper, our model provides 65 frames-per-second on 16-frames input clips. Results on UCF-Sports and JHMDB-21 datasets present comparable accuracy to State-of-the-Art approaches. We obtain F-mAP of 90.39% and 76.29% with gains of 2.19% and 1.89%, respectively. On the JHMDB-21dataset, V-mAP reaches 90.7%, 90.0% and 69.5% with gains of 2.9%, 4.3% and 11.4% at IoU thresholds of 0.2, 0.5, and 0.75.
Keywords:
Feature extraction
Three-dimensional displays
Licenses
Data mining
Location awareness
Spatiotemporal phenomena
Shape
Action localization
real-time performance
deep learning

Journal

IEEE Access cover
IEEE Access
IF:
3.6
Papers:
9.8W
Citations:
29.4W

Organization

No organization information available