arrow
返回

DARE: Deceiving Audio-Visual speech Recognition model

delete2021-11-01
delete12
PRE
AI
A
Anup Kumar Gupta *
P
Puneet Gupta
DOI:10.1016/j.knosys.2021.107503delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Audio-Visual speech recognition (AVSR) is an effective way to predict text corresponding to the spoken words using both audio and face videos, even in a noisy environment. These models find extensive applications in various fields like assisting hearing-impaired, biometric verification and speaker verification. Adversarial examples are created by adding imperceptible perturbations to the original input resulting in an incorrect classification by the deep learning models. Attacking an AVSR model is quite challenging, as both audio and visual modalities complement each other. Moreover, the correlation between audio and video features decreases while crafting an adversarial example, which can be used for detecting the adversarial example. We propose an end-to-end targeted attack, Deceiving Audio-visual speech Recognition model (DARE), which successfully performs an imperceptible adversarial attack while remaining undetected by the existing synchronisation-based detection network, SyncNet. To this end, we are the first to perform an adversarial attack that fools the AVSR model and SyncNet simultaneously. Experimental results on the publicly available dataset using state-of-the-art AVSR model reveal that the proposed attack can successfully deceive the AVSR model while remaining undetected. Furthermore, our DARE attack circumvents the well-known defences while maintaining a 100% targeted attack success rate. (C) 2021 Elsevier B.V. All rights reserved.
Keyword:
Audio-Visual Speech Recognition
Adversarial attacks
Cross-modality
Detection network

期刊

K
Knowledge-Based Systems
IF:
7.6
论文数:
1.3W
被引数:
4.5W

机构

I
indian institute of technology system (iit system)
学者数:
9.5W
论文数: 9.9W
被引数: 93
引用论文

引用论文

The Eye-RIS CMOS Vision System
err2008-01-01
err0
PREAI
errÁngel Rodríguez-Vázquez; Rafael Domínguez-Castro; Francisco Jiménez-Garrido; Sergio Morillas; Juan Listán; Luis Alba; Cayetana Utrera; Servando Espejo; Rafael Romay
err分享
err收藏
err分享
err收藏
Robust Audio-Visual Speech Recognition Under Noisy Audio-Video Conditions
err2014-02-01
err54
errOAAI
errStewart, Darryl; Seymour, Rowan; Pass, Adrian; Ming, Ji
err分享
err收藏
A low-query black-box adversarial attack based on transferability一种基于可传递性的低查询黑盒对抗攻击
err2021-08-01
err15
PREAI
errDing, Kangyi; Liu, Xiaolei; Niu, Weina; Hu, Teng; Wang, Yanping; Zhang, Xiaosong
err分享
err收藏
err分享
err收藏
Audio-visual event recognition in surveillance video sequences
err2007-02-01
err123
PREAI
errCristani, Marco; Bicego, Manuele; Murino, Vittorio
err分享
err收藏
err分享
err收藏
Automatic classification of personal video recordings based on audiovisual features基于视听特征的个人录像自动分类
err2015-11-01
err4
PREAI
errBarbancho, Ana M.; Tardon, Lorenzo J.; Lopez-Carrasco, Javier; Eggink, Jana; Barbancho, Isabel
err分享
err收藏
Recent advances in the automatic recognition of audiovisual speech
err2003-09-01
err486
PREAI
errPotamianos, G; Neti, C; Gravier, G; Garg, A; Senior, AW
err分享
err收藏
学者 查看更多内容