Return
Reconstruct the Pruned Model Without Retraining
DOI:10.1109/JSTSP.2025.3568224.png)
Abstract
En 中文
Structured pruning is a promising hardware-friendly compression technique for large language models(LLMs), which is expected to be retraining-free to avoid the enormous retraining cost. This retraining-free paradigm involves pruning criteria to define the architecture and distortion reconstruction to restore performance. However, existing reconstruction algorithms often exhibit limited generalizability and can lead to significant error accumulation. To address them, we propose a Linear Interpolation-based Adaptive Reconstruction (LIAR) framework. By applying linear interpolation to the preserved weights, LIAR minimizes the accumulated error and achieves efficient and accurate reconstruction. Furthermore, LIAR is compatible with diverse pruning criteria and modules. Our evaluations on GLUE, SQuAD, WikiText, and reasoning benchmarks show that LIAR enables a BERT model to maintain 98% accuracy even after removing 50% of its parameters and achieves 2.56× performance enhancement for LLaMA-7B under the 50% pruning ratio within 1 minute.
Keywords:
Structured pruning
retraining-free compression
distortion reconstruction
Journal
IF:
13.7
Papers:
1.9K
Citations:
1.1W

