arrow
Return

Relation-mining self-attention network for skeleton-based human action recognition

delete2023-07-01
delete32
PRE
AI
K
Kumie Gedamu *
Y
Yanli Ji
杨阳 (Yang Yang)
H
Heng Tao Shen
DOI:10.1016/j.patcog.2023.109455delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Modeling spatiotemporal global dependencies and dynamics of body joints are crucial to recognizing ac-tions from 3D skeleton sequences. We propose a Relation-mining Self-Attention Network (RSA-Net) for skeleton-based human action recognition. The proposed RSA-Net is motivated by two important obser-vations: (1) body joint relationships can be modeled independently as pairwise and unary to reduce the difficulty of action feature learning. (2) Computing action semantics and position information inde-pendently removes noisy correlations over heterogeneous embedding. The proposed RSA-Net contains pairwise self-attention, unary self-attention, and position embedding attention modules. The pairwise self-attention captures the relationship between every two body joints. The unary self-attention learns a general correlation features among one key joint over all other query joints. The position embedding attention module computes the correlation between action semantics and position information indepen-dently with separate projection matrices. Extensive evaluations are performed in the NTU-60, NTU-120, and UESTC datasets with CS, CV, CSet, and A-view evaluation benchmarks. The proposed RSA-Net outper-forms existing transformer-based approaches and comparable results with state-of-the-art graph ConvNet methods. The source code is available in Github1.(c) 2023 Elsevier Ltd. All rights reserved.
Keywords:
Action recognition
Relation-mining self-attention
Pairwise self-attention
Unary self-attention
Position attention

Journal

Pattern Recognition cover
Pattern Recognition
IF:
7.6
Papers:
1.3W
Citations:
4.5W

Organization

No organization information available