arrow
Return

CARL: Unsupervised Code-Based Adversarial Attacks for Programming Language Models via Reinforcement Learning

delete2024-12-28
delete0
delete
OA
AI
K
Kaichun Yao
H
Hao Wang
C
Chuan Qin
H
Hengshu Zhu
Y
Yanjun Wu
张立波 cover
张立波 (Libo Zhang) *
DOI:10.1145/3688839delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Code based adversarial attacks play a crucial role in revealing vulnerabilities of software system. Recently, pretrained programming language models (PLMs) have demonstrated remarkable success in various significant software engineering tasks, progressively transforming the paradigm of software development. Despite their impressive capabilities, these powerful models are vulnerable to adversarial attacks. Therefore, it is necessary to carefully investigate the robustness and vulnerabilities of the PLMs by means of adversarial attacks. Adversarial attacks entail imperceptible input modifications that cause target models to make incorrect predictions. Existing approaches for attacking PLMs often employ either identifier renaming or the greedy algorithm, which may yield sub-optimal performance or lead to high inference times. In response to these limitations, we propose CARL, an unsupervised black-box attack model that leverages reinforcement learning to generate imperceptible adversarial examples. Specifically, CARL comprises a programming language encoder and a perturbation prediction layer. In order to achieve more effective and efficient attack, we cast the task as a sequence decision-making process, optimizing through policy gradient with a suite of reward functions. We conduct extensive experiments to validate the effectiveness of CARL on code summarization, code translation, and code refinement tasks, covering various programming languages and PLMs. The experimental results demonstrate that CARL surpasses state-of-the-art code attack models, achieving the highest attack success rate across multiple tasks and PLMs while maintaining high attack efficiency, imperceptibility, consistency, and fluency.
Keywords:
Code Adversarial Attacks
Programming Language Models
Reinforcement Learning

Journal

A
ACM Transactions on Software Engineering and Methodology
IF:
6.2
Papers:
1.2K
Citations:
3.4K

Organization

I
institute of software, cas
Scholars:
445
Papers: 387
Citations: 0
U
university of chinese academy of sciences, cas
Scholars:
4.1W
Papers: 3.8W
Citations: 75
C
chinese academy of sciences
Scholars:
56.1W
Papers: 44.8W
Citations: 704
researcher View more organizations