arrow
Return

A singular learning theory for unified large language model pruning

delete2025-11-15
delete0
PRE
AI
王昕宇 cover
王昕宇 (Xinyu Wang)
Z
Zhaoxin Fan *
F
Faguo Wu
H
Hongwei Zheng
Y
Yuanze Hu
G
Gen Li
Z
Zhichao Yang
N
null Ye
Y
Yifan Sun *
W
Wenjun Wu
DOI:10.1016/j.neucom.2025.132044delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Pruning is essential for efficiently scaling large language models (LLMs), yet most current methods remain empirical and lack a unified theoretical foundation. In this work, we introduce a geometric framework for LLM pruning based on Singular Learning Theory (SLT). We show that the activation space of LLMs can be understood as an algebraic variety, and that pruning can be interpreted as a geometric folding operation within this space. Furthermore, we demonstrate that the real log canonical threshold (RLCT) serves as a principled metric for quantifying the information preserved after pruning. Building on these insights, we propose a theoretical framework that formulates pruning as an optimization problem balancing information retention and model complexity. Within this framework, we rigorously derive an upper bound on the information loss caused by pruning. We also show that leading methods such as Wanda and SparseGPT naturally fit as special cases of our approach.

Journal

Neurocomputing cover
Neurocomputing
IF:
6.5
Papers:
2.5W
Citations:
6.5W

Organization

B
Beihang University
Scholars:
5.2W
Papers: 4.1W
Citations: 37
R
Renmin University of China
Scholars:
8.1K
Papers: 7.7K
Citations: 1.1W
B
Beijing Academy of Blockchain and Edge Computing
Scholars:
15
Papers: 20
Citations: 0
researcher View more organizations