Return
ExMe: Keep Improving With Extrapolation and Merging
Y
B
Y
高
DOI:10.1109/tkde.2026.3708509.png)
Abstract
En 中文
Model merging, a typical parameter editing technique, combines different models in expectation of superior performance, without relying on the substantial computational resources and high-quality annotated data that instruction tuning heavily depends on. However, this approach lacks well-defined optimization objectives during parameter integration, making consistent performance gains in the merged model unattainable. This paper attempts to provide a clear optimization direction for model merging. We first validate the effectiveness of model extrapolation during the instruction fine-tuning phase, and then propose the Extrapolation Merging paradigm, which continuously improves model performance without requiring additional computational resources or data. Through the extrapolation method, we provide a clear direction for model merging and achieve local optimization region search, thereby ensuring that the merged model exhibits significant performance improvements. We conduct experiments on multiple different tasks, and the results show that our method consistently and stably enhances model performance after fine-tuning. Our main contributions are: (1) validating the effectiveness of model extrapolation between the base model and the supervised fine-tuned model; (2) revealing that extrapolation has a performance boundary where moderate coefficients improve performance but excessive ones cause degradation; and (3) proposing ExMe, which combines interpolation and extrapolation to achieve stable performance gains.
Keywords:
Large language model
model merging
model extrapolation
Journal
IF:
10.4
Papers:
6.7K
Citations:
3.2W
