Return
Real-Time 3-D Human Action Recognition Based on Hyperpoint Sequence
DOI:10.1109/TII.2022.3223225.png)
Abstract
En 中文
Real-time 3-D human action recognition has broad industrial applications, such as surveillance, human-computer interaction, and healthcare monitoring. By relying on complex spatio-temporal local encoding, most existing point cloud sequence networks capture spatio-temporal local structures to recognize 3-D human actions. To simplify the point cloud sequence modeling task, we propose a lightweight and effective point cloud sequence network referred to as SequentialPointNet for real-time 3-D action recognition. Instead of capturing spatio-temporal local structures, SequentialPointNet encodes the temporal evolution of static appearances to recognize human actions. First, we define a novel type of point data, hyperpoint, to better describe the temporally changing human appearances. A theoretical foundation is provided to clarify the information equivalence property for converting point cloud sequences into hyperpoint sequences. Second, the point cloud sequence modeling task is decomposed into a hyperpoint embedding task and a hyperpoint sequence modeling task. Specifically, for hyperpoint embedding, the static point cloud technology is employed to convert point cloud sequences into hyperpoint sequences, which introduces inherent frame-level parallelism; for hyperpoint sequence modeling, a hyperpoint-mixer module is designed as the basic building block to learning the spatio-temporal features of human actions. Extensive experiments on three widely-used 3-D action recognition datasets demonstrate that the proposed SequentialPointNet achieves a competitive classification performance with up to 10x faster than existing approaches.
Keywords:
3-D action recognition
hyperpoint
point cloud sequence
SequentialPointNet
Journal
IF:
9.9
Papers:
8.3K
Citations:
6.0W

