arrow
Return

Mitigating sensitive information leakage in LLMs4Code through machine unlearning

delete2026-01-22
delete0
PRE
AI
S
Shanzhi Gu
Z
Zhaoyang Qu
R
Ruotong Geng
M
Mingyang Geng
S
Shangwen Wang
C
Chuanfu Xu
H
Haotian Wang
Z
Zhipeng Lin
D
Dezun Dong
DOI:10.1016/j.neunet.2026.108606delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
• Machine unlearning is a promising way to simultaneously mitigate the privacy concerns of LLMs4Code while maintaining their code generation capabilities at the same time. Specifically, unlearning can decrease the leak rate of AIXCoder by more than 50% while only bringing a negligible side effect to code generation. • After unlearning, LLMs4Code learn to adopt diverse forms to prevent the leakage of sensitivities, in which the most popular one is to replace the sensitive fields with variable names and abbreviations. • After unlearning, LLMs4Code become more likely to leak the privacy indirectly, which means they tend to leak the information that is not explicitly queried. This suggests that future works should also take into consideration the indirect privacy leakage for a more robust unlearning process. • All code and data in this study are publicly available at https://doi.org/10.5281/zenodo.14729266 .

Journal

Neural Networks cover
Neural Networks
IF:
6.3
Papers:
7.7K
Citations:
3.0W

Organization

A
Academy of Military Sciences
Scholars:
156
Papers: 54
Citations: 0
N
national university of defense technology
Scholars:
4.1K
Papers: 1.3K
Citations: 0
B
bytedance
Scholars:
87
Papers: 39
Citations: 3
researcher View more organizations