arrow
Return

Multi-objective optimization collaborating with deep reinforcement learning: Adversarial audio attacks

delete2025-08-09
delete0
PRE
AI
P
Pengchuan Wang
W
Wen Cui
D
Deqiang Li *
Q
Qianmu Li *
DOI:10.1016/j.neucom.2025.131124delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
With the increasing prevalence of human-computer speech interaction, the security of Automatic Speech Recognition (ASR) models in Speech-to-Text (STT) systems has become a critical concern. Generating targeted and concealed adversarial speech examples poses significant challenges due to noise disturbances, inefficient querying processes, and limited consideration of speech timing characteristics. This paper introduces a novel black-box generative audio attack method, Deep Q-Network-Driven Multi-objective Optimization in Adversarial Audio Attack (DQMOA). For the first time, deep reinforcement learning is integrated into adversarial audio attacks. By leveraging deep Q-network capabilities to map solution spaces to attack behaviors, along with a reward and punishment mechanism and experience pool, DQMOA delivers optimal decision-making for multi-objective genetic algorithm. This approach substantially reduces query counts, achieves high attack success rates, and ensures low word error rates with superior speech naturalness. Extensive evaluations are conducted on the Mozilla Common Voice and LibriSpeech datasets across three commercial ASR platforms, where DQMOA consistently outperforms five state-of-the-art baselines. Further experiments on Whisper large-v2/v3 and real-world over-the-air scenarios demonstrate its strong generalization ability and practical applicability. DQMOA achieves high attack success rates with over 10 % fewer queries, while preserving low word error rates and high perceptual audio quality. These findings highlight the potential of reinforcement learning-guided evolutionary optimization for robust and stealthy adversarial audio attack generation.
Keywords:
adversarial audio attack
deep reinforcement learning
black-box attack
automatic speech recognition
multi-objective optimization

Journal

Neurocomputing cover
Neurocomputing
IF:
6.5
Papers:
2.5W
Citations:
6.5W

Organization

N
Nanjing University of Science and Technology
Scholars:
5.6K
Papers: 2.2K
Citations: 25
N
nanjing university of posts and telecommunications
Scholars:
3.4K
Papers: 1.4K
Citations: 0