arrow
Return

Adaptive Pruning for Large Language Models With Structural Importance Awareness

delete2026-01-23
delete0
PRE
AI
H
Haotian Zheng
J
Jinke Ren
Y
Yatong Han
Y
Yushan Sun
R
Ruichen Zhang
W
W. R. Zhang
Z
Z. Merrick Li
D
Dusit Niyato
S
Shuguang Cui
DOI:10.1109/JIOT.2026.3654102delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
The recent advancements in large language models (LLMs) have significantly enhanced language understanding and content generation capabilities. However, the deployment of LLMs on resource-constrained Internet of Things (IoT) devices remains challenging due to their substantial computational and storage requirements. To address this issue, we propose a novel LLM pruning method, termed structurally-aware adaptive pruning (SAAP), to reduce computational and storage costs for LLMs while maintaining model performance. Specifically, SAAP first leverages maximum likelihood estimation to calibrate traditional structural importance metrics for LLM pruning. Next, it employs a Bayesian fusion approach to address the predictive uncertainty in multigranularity metrics, enabling accurate assessments of structural importance for LLMs. Then, SAAP introduces a cross-layer importance alignment mechanism based on quantile mapping, which normalizes layer-wise importance scores to ensure consistent pruning from a global perspective. Furthermore, SAAP develops an efficient block-wise fine-tuning strategy for enhancing the performance of the LLM after pruning. To validate the effectiveness of SAAP, we conduct extensive experiments on nine open-source LLMs across two representative tasks—language modeling and zero-shot classification. Experimental results show that SAAP consistently outperforms several baseline methods, achieving accuracy improvements of 2.5%, 2.63%, and 2.44% on LLaMA-7B, Vicuna-7B, and LLaMA-13B when the pruning ratio is 50%. Finally, SAAP is implemented on a testbed—NVIDIA Jetson AGX Orin 32GB Developer Kit. Test results demonstrate that compared to the foundation LLM, SAAP enhances the inference speed by 86.86% at a pruning ratio of 50%, highlighting its potential for practical deployment on resource-constrained IoT devices.
Keywords:
Fine-tuning
large language model (LLM)
model pruning
structural importance

Journal

IEEE Internet of Things Journal cover
IEEE Internet of Things Journal
IF:
8.9
Papers:
1.4W
Citations:
7.8W

Organization

T
the chinese university of hong kong
Scholars:
4.2K
Papers: 2.0K
Citations: 0
C
chinese university of hong kong
Scholars:
2.4K
Papers: 1.2K
Citations: 0
N
nanyang technological university
Scholars:
2.5K
Papers: 1.6K
Citations: 1
H
harbin engineering university
Scholars:
5.4K
Papers: 1.9K
Citations: 0
researcher View more organizations