arrow
Return

OOPS: Outlier-aware and quadratic programming based structured pruning for large language models

delete2025-11-25
delete0
PRE
AI
J
Jiateng Wei
S
Siqi Li
J
Jingyang Xiang
J
Jun Chen
X
Xiaobin Wei
Y
Yunliang Jiang
Y
Yong Liu
DOI:10.1016/j.neunet.2025.108332delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
The large model size and resource consumption of Large Language Models (LLMs) limit their deployment and application in many scenarios. Structured pruning offers a solution to this challenge. Based on the need for retraining after pruning, structured pruning methods for LLMs fall into two categories: retraining-free and retraining-based. Retraining-free methods often result in significant performance degradation, while retraining-based methods may require substantial computational resources. To address these limitations, we propose a structured pruning framework named OOPS (Outlier-aware and quadratic prOgramming based Structured Pruning). It comprises three key components: outlier-aware pruning unit selection, quadratic programming based reconstruction, and layer-wise distillation. By employing the first two components, OOPS prunes models without the requirement of retraining, outperforming existing retraining-free methods. When further incorporating layer-wise distillation to train the pruned layers individually, OOPS surpasses other retraining-based methods with lower computational costs. We evaluate the effectiveness of OOPS on 11 models from 4 LLM families across multiple tasks, demonstrating its superior performance compared to state-of-the-art methods in both retraining-free and retraining-based settings.

Journal

Neural Networks cover
Neural Networks
IF:
6.3
Papers:
7.8K
Citations:
3.0W

Organization

W
wasu media & network co.,ltd
Scholars:
1
Papers: 1
Citations: 0
Z
Zhejiang Normal University
Scholars:
1.3W
Papers: 8.4K
Citations: 1.2W
Z
zhejiang university
Scholars:
17.6W
Papers: 12.1W
Citations: 152
researcher View more organizations