Return
Encoded motion image-based dynamic hand gesture recognition
DOI:10.1007/s00371-021-02259-3.png)
Abstract
En 中文
Dynamic hand gesture recognition is a crucial need in a smart human-computer interaction (HCI) system. Dynamic imaging has been recently introduced as a gesture description paradigm for simultaneously capturing spatial, temporal, and structural information from the depth video. However, existing techniques based on dynamic images cannot differentiate gesture movements that follow the same path but in opposite directions, for example, moving a hand down versus moving a hand up. To solve the issue, we have proposed an approach in which a gesture depth video is converted into a single image called an encoded motion image (EMI). The EMI has been given to a modified pre-trained 2D-CNN(two-dimensional convolutional neural network) based on VGG-19 to classify gestures present in the depth video. The experiments were carried out on two datasets: a multi-modal large-scale EgoGesture and MSR Gesture 3D datasets. For the EgoGesture dataset, the proposed method achieved an accuracy of 90.63%. Such a result provides state-of-the-art accuracy when employing this large-scale dataset of 83 classes and the 2D-CNN approach. For the MSR Gesture 3D dataset, the proposed method accuracy is 99.24%, which outperforms the state-of-the-art methods. This work also highlights the recognition accuracy and precision of each gesture. Instead of high-end systems like GPU, the experiments are conducted using a web-based data science environment called Kaggle to demonstrate the work's economic efficiency.
Keywords:
Dynamic hand gesture
Human-computer interaction
Dynamic imaging
2D-CNN
AI Summary
Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.
Journal
IF:
2.9
Papers:
4.6K
Citations:
6.5K

