Return
3D Skeleton-Based Non-Autoregressive Human Motion Prediction Using Encoder-Decoder Attention-Based Model
DOI:10.1109/TETCI.2024.3418828.png)
Abstract
En 中文
An encoder-decoder attention-based model has been employed to predict human action using a 3D skeleton-based human activity dataset. It offers and advocates a non-autoregressive approach to leverage human motion prediction effectively. The encoder framework employs a spatiotemporal attention mechanism to capture spatiotemporal features. While the decoder uses a motion attention mechanism to predict human motion. An extensive experiment has been carried out with two datasets, Human3.6 M and AMASS, to validate the performance of the proposed model. The proposed model outperforms the best available SOTA methods on H3.6 M by 13.63%, 14.28%, 19.14%, and 12.87% with 80 ms, 160 ms, 320 ms, and 400 ms, respectively, in terms of average Euler angle error. Similar results are observed with the AMASS dataset. The proposed model predicts human motion more accurately than the previous best methods by 2.44%, 1%, 1.25% and 1.3% in terms of Euler angle error, joint angle error, positional error, & percentage of correct key point, respectively. Furthermore, a separate discussion has been carried out to analyse the effect of autoregressive & non-autoregressive approaches for human motion prediction.
Keywords:
Transformers
Predictive models
Decoding
Data models
Vectors
Training
Task analysis
Non-autoregressive
encoder-decoder attention
motion transformer
spatiotemporal attention
human skeleton
human motion prediction
Journal
I
IF:
6.5
Papers:
1.4K
Citations:
4.5K

