Return
An Overview of Low-Rank Structures in the Training and Adaptation of Large Models: From Implicit Low-Dimensionality to Efficient Training and Fine-Tuning [Special Issue on the Mathematics of Deep Learning]
L
T
B
S
Q
P
Z
C
DOI:10.1109/MSP.2026.3666749.png)
Abstract
En 中文
The advent of deep learning has immeasurably changed the ways we process data in signal processing and machine learning. However, training and deploying modern deep learning models demand substantial computational resources, raising concerns about exorbitant training costs, GPU shortages, and heightened energy consumption. Several lines of research over the last decade have explored the emergence of low-dimensional structures during the training process, where basic elements such as weight matrices and representations tend to be approximately low rank even though not explicitly trained to be. These low-dimensional structures arise in part due to the implicit bias of the methods used to train deep networks, providing the potential to partially explain why deep models need fewer samples than the number of model parameters. This implicit low dimensionality has inspired the exploration of low-rank structures in training and fine-tuning large-scale deep learning models more efficiently. In this article, we review recent exciting advances in using low-rank structure in deep learning and aim to clarify the mathematical foundations underlying their design. Specifically, we highlight key insights from a rich line of research focused on theoretically understanding and leveraging low-rank structures in deep learning both at every iteration of training as well as at the global minimum.
Keywords:
Deep learning
Data processing
Signal processing
Large language models
Ranking (statistics)
Training
Machine learning
Resource management
Journal
IF:
9.6
Papers:
1.1W
Citations:
1.7W
