Return
Mitigating sensitive information leakage in LLMs4Code through machine unlearning
DOI:10.1016/j.neunet.2026.108606.png)
Abstract
En 中文
• Machine unlearning is a promising way to simultaneously mitigate the privacy concerns of LLMs4Code while maintaining their code generation capabilities at the same time. Specifically, unlearning can decrease the leak rate of AIXCoder by more than 50% while only bringing a negligible side effect to code generation. • After unlearning, LLMs4Code learn to adopt diverse forms to prevent the leakage of sensitivities, in which the most popular one is to replace the sensitive fields with variable names and abbreviations. • After unlearning, LLMs4Code become more likely to leak the privacy indirectly, which means they tend to leak the information that is not explicitly queried. This suggests that future works should also take into consideration the indirect privacy leakage for a more robust unlearning process. • All code and data in this study are publicly available at https://doi.org/10.5281/zenodo.14729266 .
Journal
IF:
6.3
Papers:
7.7K
Citations:
3.0W

