arrow
Return

SampleLLM: Prune LLMs via learnable structure sampling

delete2026-06-05
delete0
PRE
AI
J
Jiateng Wei
S
Siqi Li
X
Xin Yuan
J
Jingyang Xiang
L
Longqi Wang
C
Chuang Guo
J
Jun Chen *
B
Baochang Zhu
J
Jie Ren
W
Weiwei Liu
Y
Yong Liu *
DOI:10.1016/j.neucom.2026.134215delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
As the scale of large language models (LLMs) continues to grow, there is an increasing need to reduce computational overhead. Structured pruning has been proven to be an effective technique for model compression. However, existing methods are often hindered by a reliance on manually defined heuristic metrics and a fragmented “prune-then-retrain” pipeline that ignores the correlations between different pruning structures. To address these drawbacks, we introduce SampleLLM, which formulates pruning as a learnable sampling process. By jointly optimizing sampling parameters and LoRA weights, SampleLLM captures dynamic changes in the importance of pruning structures and eliminates the need for manual pruning metrics. We also introduce a self-distillation mechanism to further align the pruned model with its dense counterpart. Extensive experiments on 9 models across 5 families demonstrate that SampleLLM outperforms state-of-the-art structured pruning methods. For example, our 40% pruned LLaMA2-70B model retains 98% of its original performance. Moreover, at the same compression rate, it achieves superior performance and inference speedup compared to 2:4 semi-structured pruning schemes.

Journal

Neurocomputing cover
Neurocomputing
IF:
6.5
Papers:
2.5W
Citations:
6.5W

Organization

Z
zhejiang normal university
Scholars:
3.0K
Papers: 1.1K
Citations: 0
Z
zhejiang university
Scholars:
17.6W
Papers: 12.1W
Citations: 152