Return
ETV-Attack: Efficient text-driven visual-variable adversarial attacks on visual question answering with pre-trained language models
DOI:10.1016/j.patcog.2026.113202.png)
Abstract
En 中文
• This work provides an overall evaluation of LLM-based VQA in attack. • This work designs a text-driven embedding-based attack method. • This work proposes a text-driven augmentation module for VQA. • This work achieves new SOTA results on relevant benchmarks.
Keywords:
LLM-based VQA
text-driven attack
visual-variable adversarial attacks
pre-trained language models
VQA augmentation
Journal
IF:
7.6
Papers:
1.3W
Citations:
4.5W

