arrow
Return

DEFT: Data-Efficient Fine-Tuning Through Multi-Dimensional Data Selection

delete2026-01-01
delete0
PRE
AI
S
Shaojie Dai
X
Xin Liu
余月 cover
余月 (Yue Yu) *
DOI:10.1109/TASLPRO.2025.3642562delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Instruction tuning has emerged as a predominant method for adapting large language models (LLMs) to downstream tasks, with prevailing approaches predominantly relying on scaling up instruction data to enhance model performance. However, growing evidence suggests that indiscriminate data scaling may yield suboptimal results, as the absence of systematic evaluation criteria often leads to redundant or low-quality samples in instruction datasets. Consequently, in this paper, we propose DEFT, a multi-dimensional data selection framework that assesses instruction data from four perspectives: complexity, quality, knowledge and diversity. For complexity and quality, we develop Evol-Ranking to distill ranking capabilities from teacher models (e.g., gpt-3.5-turbo) to specialized student models. Furthermore, we propose refinement distillation to progressively optimize the student model. For knowledge, we define the average negative log-probability of text on a given LLM as knowledge, providing model-aware measurement. For diversity, we first obtain semantic representation of each sample, then calculate the similarity between samples. Finally, we ensemble all dimensions mentioned above through an ensemble scoring mechanism to select the data for instruction fine-tuning. Extensive experiments performed on MT-Bench and AlpacaEval demonstrate that DEFT performs better or on pair with the state-of-the-art open-source alignment models with only 6,000 SFT training samples.
Keywords:
Data models
Complexity theory
Training data
Training
Speech processing
Semantics
Adaptation models
Tuning
Measurement
Large language models
Data-efficient fine-tuning
large language models (LLMs)
instruction fine-tuning
data selection

Journal

I
IEEE Transactions on Audio Speech and Language Processing
IF:
0
Papers:
151
Citations:
0

Organization

I
institute of computing technology, cas
Scholars:
1.0K
Papers: 877
Citations: 1
C
chinese academy of sciences
Scholars:
55.9W
Papers: 44.7W
Citations: 704