Return
CARD: Contrastive Adversarial Representation Distillation for Robustness of Lightweight Large Language Models
X
W
H
D
洪
X
罗
DOI:10.1109/mnet.2026.3662210.png)
Abstract
En 中文
In this paper, we propose a Contrastive Adversarial Representation Distillation (CARD) framework to address the adversarial robustness degradation in lightweight large language models (LLMs). Existing adversarial knowledge distillation (AKD) methods merely transfer robustness by mimicking the teacher’s output logits, failing to capture deeper defensive knowledge. Our CARD framework innovatively designs a contrastive representation distillation loss, which aligns feature representations between teacher and student models for clean- adversarial sample pairs, enabling representation invariance to be directly inherited as a core defense mechanism. Experiments on IMDB sentiment classification demonstrate that CARD achieves a 58.33% averaged Attack Failure Rate (Avg AFR) compared to 28.33% for state-of-the-art methods, with minimal clean accuracy degradation (<inline-formula xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"> <tex-math notation="LaTeX">$\leq 1$ </tex-math></inline-formula>%). This work establishes representation invariance distillation as a superior paradigm for enhancing adversarial robustness in lightweight LLMs<xref ref-type="fn" rid="fn1" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</xref>.
Keywords:
Robustness
Training
Computational modeling
Perturbation methods
Contrastive learning
Data models
Text categorization
Standards
Security
Large language models
Adversarial machine learning
Contrastive learning
Journal
IF:
6.3
Papers:
2.6K
Citations:
1.1W
