arrow
Return

Memory-efficient neural network training via gradient compression through continuous basis tracking

delete2026-03-07
delete0
PRE
AI
N
null Meng
M
Muhammad A.A. Abdelgawad
P
Peng Jing
R
Ray C.C. CHEUNG
H
Hong Yan
DOI:10.1016/j.neucom.2026.133278delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Large Language Models (LLMs) have achieved unprecedented success across a wide range of tasks, but their training and fine-tuning require significant computational resources and memory due to the vast number of parameters and optimizer states. While the recent method GaLore leveraged low-rank gradient projection to reduce memory consumption, it suffered from limited expressiveness or high computational overhead, mainly due to frequent full-rank singular value decompositions (SVD). In this work, we propose InGaLore, a novel memory- and time-efficient neural network training framework that introduces truncated incremental SVD to update the gradient projection basis continuously. Our approach tracks the evolving subspace of gradients, enabling efficient compression in low-rank gradients and reduced computational cost. We further analyze and address the loss spike phenomenon observed during projection matrix updates, providing a new explanation and an effective mitigation strategy. Experiments on both pre-training and fine-tuning tasks with LLMs demonstrate that InGaLore and ERInGaLore deliver superior or comparable accuracy to existing methods, while significantly reducing wall-clock time by 15% and maintaining low memory usage. Importantly, the wall-clock time savings provided by our method become increasingly significant as model scale grows.
Keywords:
gradient compression
low-rank approximation
incremental SVD
neural network training
memory efficiency

Journal

Neurocomputing cover
Neurocomputing
IF:
6.5
Papers:
2.5W
Citations:
6.5W

Organization

No organization information available
Cited Papers

Cited Papers

No cited papers available