arrow
Return

Difference-guided multi-scale spatial-temporal representation for sign language recognition

delete2023-07-30
delete3
PRE
AI
L
Liqing Gao
L
Lianyu Hu
F
Fan Lyu
朱磊 cover
朱磊 (Lei Zhu)
L
Liang Wan
C
Chi‐Man Pun
W
Wei Feng *
DOI:10.1007/s00371-023-02979-8delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Sign language recognition (SLR) is a challenging task, which requires a thorough understanding of spatial-temporal visual features for translating it into comprehensible written or spoken language. However, existing SLR methods ignore the importance of key spatial-temporal representation due to its sparsity and inconsistency in space and time. To solve this problem, we present a difference-guided multi-scale spatial-temporal representation (DMST) learning model for SLR. In DMST, we devise two modules: (1) key spatial-temporal representation, to extract and enhance key spatial-temporal information by a spatial-temporal difference strategy and (2) multi-scale sequence alignment, to perceive and fuse multi-scale spatial-temporal features and achieve sequence mapping. The DMST model outperforms state-of-the-art performance on four public sign language datasets, which demonstrates the superiority of DMST model and the significance of key spatial-temporal representation for SLR.
Keywords:
Sign language recognition (SLR)
Key spatial-temporal representation
Multi-scale sequence alignment

Journal

Visual Computer cover
Visual Computer
IF:
2.9
Papers:
4.6K
Citations:
6.5K

Organization

T
tianjin university
Scholars:
7.9W
Papers: 5.7W
Citations: 88
U
University of Macau
Scholars:
1.1W
Papers: 1.3W
Citations: 2.0W
researcher View more organizations