arrow
Return

Optimized Multi-Modal Conformer-Based Framework for Continuous Sign Language Recognition

delete2025-01-01
delete0
delete
OA
AI
N
Neena Aloysius
P
Prema Nedungadi
DOI:10.1109/OJCS.2025.3564828delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
This study introduces Efficient ConSignformer, a novel framework advancing Continuous Sign Language Recognition (CSLR) by optimizing the Conformer-based CSLR model, ConSignformer. Central to this advancement is the Sign Query Attention (SQA) module, a computationally efficient self-attention mechanism that enhances both performance and scalability, resulting in the Efficient Conformer. Efficient ConSignformer integrates video embeddings from dual-modal CNN pipelines that process heatmaps and RGB videos, along with temporal learning layers tailored for each modality. These embeddings are further refined using the Efficient Conformer for the fused data from two modalities. To improve recognition accuracy, we employ an innovative task-adaptive supervised pretraining strategy for Efficient Conformer on a curated dataset of continuous Indian Sign Language (ISL). This strategy enables the model to effectively capture intricate data relationships during end-to-end training. Experimental results highlight the significant contributions of the SQA module and the pretraining strategy, with our model achieving competitive performance on benchmark datasets, PHOENIX-2014 and PHOENIX-2014 T. Notably, Efficient ConSignformer excels in recognizing longer sign sequences, leveraging a computationally lightweight Conformer backbone.
Keywords:
Conformer
ConSignformer
continuous sign language recognition
optimization
supervised pretraining
task-adaptive pretraining

Journal

I
IEEE Open Journal of the Computer Society
IF:
8.2
Papers:
411
Citations:
810

Organization

A
Amrita School of Computing
Scholars:
60
Papers: 35
Citations: 0