arrow
Return

Automated Pruning Framework for Large Language Models Using Combinatorial Optimization

delete2025-07-27
delete0
delete
OA
AI
P
Patcharapol Ratsapa
K
Kundjanasith Thonglek
DOI:10.3390/ai6050096delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Currently, large language models (LLMs) have been utilized in many aspects of natural language processing. However, due to their significant size and high computational demands, large computational resources are required for deployment. In this research, we focus on the automated approach for size reduction of such a model. We propose the framework to perform the automated pruning based on combinatorial optimization. Two techniques were particularly studied, i.e., particle swarm optimization (PSO) and whale optimization algorithm (WOA). The model pruning problem was modeled as a combinatorial optimization task whose the goal is to minimize model size while maintaining model accuracy. The framework systematically explores the search space to identify the most optimal pruning configurations, removing redundant or non-contributory parameters. The two optimizations, PSO and WOA, were evaluated for their ability to efficiently navigate the search space. As a result, with PSO, the proposed framework can reduce the model size of Llama-3.1-70B by 13.44% while keeping the loss of model accuracy at 19.25%; with WOA, the model size reduction is 12.07% with 22.81% loss of model accuracy. Since accuracy degradation may occur during pruning process, the framework integrates the post-process to recover the model accuracy. After this process, the pruned model loss can reduce to 12.72% and 14.83% using PSO and WOA, respectively.
Keywords:
large language models
model pruning
combinatorial optimization
particle swarm optimization
whale optimization algorithm

Journal

A
AI
IF:
5
Papers:
991
Citations:
941

Organization

No organization information available