Return
A Unified Optimization Framework for Backdoor Attacks in Large Language Models
DOI:10.1016/j.inffus.2026.104221.png)
Abstract
En 中文
• Propose a framework for backdoor attacks in large language models. • Theoretically prove the sufficiency of partial parameter updates. • Resolve gradient conflicts between clean and adversarial objectives. • Validate higher success rates with minimal clean-task performance drop.
Keywords:
Backdoor Attacks
Large Language Models
Optimization Framework
Gradient Conflict
Parameter Updates
Journal
IF:
15.5
Papers:
4.1K
Citations:
2.7W

