Return
Design and Efficacy of a Speaker Verification Method Combining CNN and Transformer for Secure Access Control
X
X
K
L
DOI:10.3390/s26165101.png)
Abstract
En 中文
In the rapidly evolving landscape of computer and mobile applications, the demand for secure access control has become increasingly pivotal. This paper introduces a speaker verification method aimed at remotely verifying an individual’s claimed identity, so as to achieve access authorization. The primary objective is to develop a deep learning network capable of eliminating redundant and irrelevant information while learning robust deep speaker embedding descriptors that capture speaker-specific idiosyncrasies. This paper presents an AI framework combining convolutional neural network and transformer architectures, enhanced by (1) a stereoscopic attention mechanism that computes fine-grained attention weights across frequency, time, and channel dimensions with fewer parameters than existing CBAM or SE modules, and (2) a multi-feature aggregation mechanism that fuses supervised and unsupervised features at the utterance level to maximize complementary speaker information. These innovations introduce a new perspective for integrating acoustic and articulatory features, effectively addressing the challenges related to short-segment speech and cross-domain scenarios. Experimental results demonstrate that the AI framework achieves competitive performance under domain mismatch conditions. The methodology acts as a catalyst for advancing application authorization, highlighting the transformative potential of AI-driven innovations in the software engineering field.
Keywords:
speaker verification
CNN
transformer
Journal
IF:
3.5
Papers:
7.1W
Citations:
20.9W
