Return
LaRA: Layer-wise rank allocation for efficient fine-tuning of pruned large language models
Y
C
S
F
Y
A
DOI:10.1016/j.ipm.2025.104538.png)
Abstract
En 中文
• Introduce LaRA, an automated framework that restores pruned LLMs via LoRA by assigning layer-specific ranks, achieving optimal performance with minimal overhead. • Propose a Rank Allocation Model, trained with policy-gradient reinforcement learning, and a Score Model environment that evaluates rank configurations as rewards. • Build a diverse dataset of architecture-accuracy pairs spanning pruning levels, LLM families, and ranks to robustly train both models. • Extensive experiments across tasks, pruning settings, and LLMs validate LaRA’s effectiveness and generality
Journal
I
IF:
6.9
Papers:
5.2K
Citations:
1.4W
