arrow
Return

Reconstruct the Pruned Model Without Retraining

delete2025-05-12
delete0
PRE
AI
P
Pingjie Wang
Z
Ziqing Fan
S
Shengchao Hu
Z
Zhe Chen
王
王延峰 (Yanfeng Wang)
王
王钰 (Yu Wang)
DOI:10.1109/JSTSP.2025.3568224delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Structured pruning is a promising hardware-friendly compression technique for large language models(LLMs), which is expected to be retraining-free to avoid the enormous retraining cost. This retraining-free paradigm involves pruning criteria to define the architecture and distortion reconstruction to restore performance. However, existing reconstruction algorithms often exhibit limited generalizability and can lead to significant error accumulation. To address them, we propose a Linear Interpolation-based Adaptive Reconstruction (LIAR) framework. By applying linear interpolation to the preserved weights, LIAR minimizes the accumulated error and achieves efficient and accurate reconstruction. Furthermore, LIAR is compatible with diverse pruning criteria and modules. Our evaluations on GLUE, SQuAD, WikiText, and reasoning benchmarks show that LIAR enables a BERT model to maintain 98% accuracy even after removing 50% of its parameters and achieves 2.56× performance enhancement for LLaMA-7B under the 50% pruning ratio within 1 minute.
Keywords:
Structured pruning
retraining-free compression
distortion reconstruction

Journal

IEEE Journal of Selected Topics in Signal Processing cover
IEEE Journal of Selected Topics in Signal Processing
IF:
13.7
Papers:
1.9K
Citations:
1.1W

Organization

S
Shanghai Jiao Tong University
Scholars:
7.8K
Papers: 2.4K
Citations: 14.8W
Cited Papers

Cited Papers

Recursive Deep Models for Semantic Compositionality Over a Sentiment Treebank
err2013-01-01
err0
PREAI
errRichard Socher; Alex Perelygin; Jean Wu; Jason Chuang; Christopher D. Manning; Andrew Ng; Christopher Potts
errShare
errSave
The PASCAL Recognising Textual Entailment Challenge
err2006-01-01
err0
PREAI
errIdo Dagan; Oren Glickman; Bernardo Magnini
errShare
errSave
EdgeBERT: Sentence-Level Energy Optimizations for Latency-Aware Multi-Task NLP Inference
err2021-10-17
err0
errOAAI
errThierry Tambe; Coleman Hooper; Lillian Pentecost; Tianyu Jia; En-Yu Yang; Marco Donato; Victor Sanh; Paul Whatmough; Alexander M. Rush; David Brooks; Gu-Yeon Wei
errShare
errSave
errShare
errSave
Neural Network Acceptability Judgments
err2019-11-01
err401
errOAAI
errWarstadt, Alex; Singh, Amanpreet; Bowman, Samuel R.
errShare
errSave
researcher View more