arrow
Return

ETV-Attack: Efficient text-driven visual-variable adversarial attacks on visual question answering with pre-trained language models

delete2026-02-04
delete0
PRE
AI
Q
Quanxing Xu
周凌 cover
周凌 (Ling Zhou)
X
Xian Zhong
F
Feifei Zhang
J
Jinyu Tian
X
Xiaohan Yu
R
Rubing Huang
DOI:10.1016/j.patcog.2026.113202delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
• This work provides an overall evaluation of LLM-based VQA in attack. • This work designs a text-driven embedding-based attack method. • This work proposes a text-driven augmentation module for VQA. • This work achieves new SOTA results on relevant benchmarks.
Keywords:
LLM-based VQA
text-driven attack
visual-variable adversarial attacks
pre-trained language models
VQA augmentation

Journal

Pattern Recognition cover
Pattern Recognition
IF:
7.6
Papers:
1.3W
Citations:
4.5W

Organization

W
wuhan university of technology
Scholars:
7.1K
Papers: 2.1K
Citations: 0
M
macau university of science and technology
Scholars:
1.2K
Papers: 580
Citations: 0
T
tianjin university of technology
Scholars:
1.8K
Papers: 549
Citations: 0
M
Macquarie University
Scholars:
1.2W
Papers: 1.5W
Citations: 2.2W
researcher View more organizations