Return
Multimodal alignment augmentation transferable attack on vision-language pre-training models
DOI:10.1016/j.patrec.2025.03.007.png)
Abstract
En 中文
Vision-language Pre-training (VLP) models have exhibited superior performance in multimodal tasks. Recent research has revealed that VLP models are vulnerable to transfer-based attacks, where adversarial examples (AEs) generated from a local surrogate model successfully deceive black-box models. Consequently, exploring the vulnerability of VLP models against transfer-based attacks is crucial. However, existing transfer-based attacks for VLP models overlook the fact that an image should not be represented by a single text, nor should a text describe only one image, resulting in constraints on the generation of AEs and limiting adversarial transferability. To address these limitations, we propose the Multimodal Alignment Augmentation Transferable Attack (MA-Attack). Specifically, we introduce a Random Augmentation Pool for images and a Reverse Attack Augmentation for texts to establish multi-to-multi image-text pairs. MA-Attack enriches the multimodal alignment information and provides sufficient representations for AE generation, thereby enhancing the adversarial transferability. Extensive experiments conducted on image-text retrieval tasks demonstrate that MA-Attack outperforms existing state-of-the-art attacks in terms of adversarial transferability.
Keywords:
Adversarial example
Vision-language pre-training model
Model vulnerability
Transfer-based attack

