arrow
Return

Parameter-Efficient Fine-Tuning for Continual Learning: A Neural Tangent Kernel Perspective

delete2026-03-06
delete0
PRE
AI
J
Jingren Liu
冀中 cover
冀中 (Zhong Ji)
Y
Yunlong Yu
J
Jiale Cao
Y
Yanwei Pang
韩军功 (Jungong Han)
X
Xuelong Li
DOI:10.1109/TPAMI.2026.3670952delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Parameter-efficient fine-tuning for continual learning (PEFT-CL) has shown promise in adapting pre-trained models to sequential tasks while mitigating catastrophic forgetting problem. However, understanding the mechanisms that dictate continual performance in this paradigm remains elusive. To unravel this mystery, we undertake a rigorous analysis of PEFT-CL dynamics to derive relevant metrics for continual scenarios using Neural Tangent Kernel (NTK) theory. With the aid of NTK as a mathematical analysis tool, we recast the challenge of test-time forgetting into the quantifiable generalization gaps during training, identifying three key factors that influence these gaps and the performance of PEFT-CL: training sample size, task-level feature orthogonality, and regularization. To address these challenges, we introduce NTK-CL, a novel framework that eliminates task-specific parameter storage while adaptively generating task-relevant features. Aligning with theoretical guidance, NTK-CL triples the feature representation of each sample, theoretically and empirically reducing the magnitude of both task-interplay and task-specific generalization gaps. Grounded in NTK analysis, our framework imposes an adaptive exponential moving average mechanism and constraints on task-level feature orthogonality, maintaining intra-task NTK forms while attenuating inter-task NTK forms. Ultimately, by fine-tuning optimizable parameters with appropriate regularization, NTK-CL achieves state-of-the-art performance on established PEFT-CL benchmarks. This work provides a theoretical foundation for understanding and improving PEFT-CL models, offering insights into the interplay between feature representation, task orthogonality, and generalization, contributing to the development of more efficient continual learning systems.
Keywords:
Parameter-efficient fine-tuning
continual learning
neural tangent Kernel
model generalization

Journal

IEEE Transactions on Pattern Analysis and Machine Intelligence cover
IEEE Transactions on Pattern Analysis and Machine Intelligence
IF:
18.6
Papers:
831
Citations:
9.8W

Organization

U
university of sheffield
Scholars:
3.4K
Papers: 1.7K
Citations: 1
T
tianjin university
Scholars:
8.0W
Papers: 5.7W
Citations: 88
C
China Telecom
Scholars:
71
Papers: 62
Citations: 121
Z
zhejiang university
Scholars:
17.6W
Papers: 12.1W
Citations: 152
researcher View more organizations