Return
Syntax-Oriented Shortcut: A Syntax Level Perturbing Algorithm for Preventing Text Data From Being Learned
DOI:10.1109/TNNLS.2025.3609842.png)
Abstract
En 中文
The vast availability of free data has been critical to the success of large language models (LLMs). With the widespread use of LLMs, more and more concerns have been raised about the unauthorized use of publicly available data. To protect data from unauthorized use for training models, researchers have proposed adding imperceptible perturbations into image data so that models would be misled by the generated shortcut features and cannot mine information from these images. However, due to the inherent discrete property and semantic complexity of texts, directly applying these methods to text will cause semantic changes, resulting in meaningless shortcut features being constructed. To tackle this problem, in this article, we design a novel Unlearnable text examples generation algorithm via syntax-oriented shortcut (UTE-SS) by incorporating the syntactic structure of texts. Specifically, we propose a syntax template generator (STG) to generate the optimal perturbing syntax for a given category, which will realize imperceptible perturbations. Then, a perturbing text generator (PTG) is designed to perturb the in-class texts with the selected syntax template to stably deviate from the original texts. Along this line, models will be misled to learn the shortcut between the syntax template and the category, so as to keep text examples unlearnable. Extensive experiments over eight advanced Transformer-based pretrained language models (PLMs) on four different natural language processing (NLP) tasks demonstrate the effectiveness and flexibility of our proposed algorithm. Our method is easy to implement, and the code is publicly available at <uri xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">https://github.com/libolb/UTE-SS.</uri>
Keywords:
Data security
privacy protection
shortcut learning
unlearnable examples
Journal
IF:
8.9
Papers:
7.5K
Citations:
7.2W

