Return
Model Inversion Attack Against Transfer Learning: Inverting a Model Without Querying It
DOI:10.1109/TDSC.2025.3548119.png)
Abstract
En 中文
Transfer learning is an important approach that produces pre-trained teacher models which can be used to quickly build specialized student models. However, recent research on transfer learning has found that it is vulnerable to various attacks, e.g., misclassification and backdoor attacks. However, it is still not clear whether transfer learning is vulnerable to model inversion attacks. Launching a model inversion attack against transfer learning scheme is challenging. Not only does the student model hide its structural parameters, but it is also not queried to the adversary. Hence, when targeting a student model, existing model inversion attacks fail, as they typically rely on querying the target model. In this paper, we initiate research into model inversion attacks against transfer learning with two novel attack methods. Both are black-box attacks, suiting different situations, that do not rely on queries to the target student model. In the first method, the adversary has the data samples that share the same distribution as the training set of the teacher model. In the second method, the adversary does not have any such samples. Experiments show that highly recognizable data records can be inverted with both of these methods. This research underscores the critical insight that even when a model is shielded from public queries, it can still be susceptible to model inversion attacks.
Keywords:
Transfer learning
model inversion attack
Journal
IF:
7.5
Papers:
2.4K
Citations:
9.6K

