Return
A singular learning theory for unified large language model pruning
DOI:10.1016/j.neucom.2025.132044.png)
Abstract
En 中文
Pruning is essential for efficiently scaling large language models (LLMs), yet most current methods remain empirical and lack a unified theoretical foundation. In this work, we introduce a geometric framework for LLM pruning based on Singular Learning Theory (SLT). We show that the activation space of LLMs can be understood as an algebraic variety, and that pruning can be interpreted as a geometric folding operation within this space. Furthermore, we demonstrate that the real log canonical threshold (RLCT) serves as a principled metric for quantifying the information preserved after pruning. Building on these insights, we propose a theoretical framework that formulates pruning as an optimization problem balancing information retention and model complexity. Within this framework, we rigorously derive an upper bound on the information loss caused by pruning. We also show that leading methods such as Wanda and SparseGPT naturally fit as special cases of our approach.
Journal
IF:
6.5
Papers:
2.5W
Citations:
6.5W

