Return
PP-LLMs: A progressive pruning approach with medium-granularity for large language models
K
Y
B
J
F
L
Y
DOI:10.1016/j.eswa.2026.133770.png)
Abstract
En 中文
• We propose a dual-metric two-stage pipeline tailored to each layer’s function. • We introduce a method quantifying layer importance for performance-aware pruning. • We design a low-cost pruning method, PP-LLMs, for effective model deployment.
Keywords:
Large language models
Model compression
Structure pruning
Layer importance metric
Journal
IF:
7.5
Papers:
2.9W
Citations:
10.2W
