1
Return

LaRA: Layer-wise rank allocation for efficient fine-tuning of pruned large language models

delete2025-12-11
delete0
PRE
AI
Y
Yuhua Zhou
C
Changhai Zhou
S
Shiyang Zhang
F
Fei Yang
Y
Yi Zhang
A
Aimin Pan
DOI:10.1016/j.ipm.2025.104538delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
• Introduce LaRA, an automated framework that restores pruned LLMs via LoRA by assigning layer-specific ranks, achieving optimal performance with minimal overhead. • Propose a Rank Allocation Model, trained with policy-gradient reinforcement learning, and a Score Model environment that evaluates rank configurations as rewards. • Build a diverse dataset of architecture-accuracy pairs spanning pruning levels, LLM families, and ranks to robustly train both models. • Extensive experiments across tasks, pruning settings, and LLMs validate LaRA’s effectiveness and generality

Journal

I
Information Processing and Management
IF:
6.9
Papers:
5.2K
Citations:
1.4W

Organization

Z
Zhejiang Lab
Scholars:
335
Papers: 165
Citations: 8.0K
C
Columbia University
Scholars:
7.1W
Papers: 6.4W
Citations: 263
F
fudan university
Scholars:
11.3W
Papers: 7.6W
Citations: 121
Z
zhejiang university
Scholars:
17.0W
Papers: 11.9W
Citations: 152
Cited Papers

Cited Papers

Citing Papers

Citing Papers