arrow
Return

Gradient Projection for Continual Parameter-Efficient Tuning

delete
delete0
delete
OA
AI
J
Jingyang Qiao
Z
Zhizhong Zhang
X
Xin Tan
Y
Yanyun Qu
张文胜 (Wensheng Zhang)
Z
Zhi Han
谢源 (Yuan Xie)
DOI:10.1109/TPAMI.2025.3587032delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Parameter-efficient tunings (PETs) have demonstrated impressive performance and promising perspectives in training large models, while they are still confronted with a common problem: the trade-off between learning new content and protecting old knowledge, leading to zero-shot generalization collapse, and cross-modal hallucination. In this paper, we reformulate Adapter, LoRA, Prefix-tuning, and Prompt-tuning from the perspective of gradient projection, and first propose a unified framework called <underline xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"><b>P</b></u>arameter <underline xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"><b>E</b></u>fficient <underline xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"><b>G</b></u>radient <underline xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"><b>P</b></u>rojection (PEGP). We introduce orthogonal gradient projection into different PET paradigms and theoretically demonstrate that the orthogonal condition for the gradient can effectively resist forgetting even for large-scale models. It therefore modifies the gradient towards the direction that has less impact on the old feature space, with less extra memory space and training time. We extensively evaluate our method with different backbones, including ViT and CLIP, on diverse datasets, and experiments comprehensively demonstrate its efficiency in reducing forgetting in class, online class, domain, task, and multi-modality continual settings.
Keywords:
Continual learning
parameter-efficient tuning
anti-catastrophic forgetting
orthogonal gradient projection
multi-modality learning

Journal

IEEE Transactions on Pattern Analysis and Machine Intelligence cover
IEEE Transactions on Pattern Analysis and Machine Intelligence
IF:
18.6
Papers:
831
Citations:
9.8W

Organization

S
Shenyang Institute of Automation
Scholars:
213
Papers: 88
Citations: 0
I
Institute of Automation
Scholars:
529
Papers: 278
Citations: 220
E
east china normal university
Scholars:
3.0W
Papers: 2.1W
Citations: 25
X
xiamen university
Scholars:
5.8W
Papers: 3.8W
Citations: 67
researcher View more organizations