arrow
返回

DynImpt: A Dynamic Data Selection Method for Improving Model Training Efficiency

delete2025-01-01
delete0
PRE
AI
W
Wei Huang
张云霄 封面图
张云霄 (Yunxiao Zhang)
S
Shangmin Guo
Y
Yu-Ming Shang *
X
Xiangling Fu *
DOI:10.1109/TKDE.2024.3482466delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Selecting key data subsets for model training is an effective way to improve training efficiency. Existing methods generally utilize a well-trained model to evaluate samples and select crucial subsets, ignoring the fact that the sample importance changes dynamically during model training, resulting in the selected subset only being critical in a specific training epoch rather than a changing training phase. To address this issue, we attempt to evaluate the significant changes in sample importance during dynamic training and propose a novel data selection method to improve model training efficiency. Specifically, the temporal changes in sample importance are considered from three perspectives: (i) loss, the difference between the predicted labels and the true labels of samples in the current training epoch; (ii) instability, the dispersion of sample importance in the recent training phase; and (iii) inconsistency, the comparison of the changing trend in the importance of an individual sample relative to the average importance of all samples in the recent training phase. Extensive experiments demonstrate that dynamic data selection can reduce computational costs and improve model training efficiency. Additionally, we find that the difficulty level of the training task influences the data selection strategy.
Keyword:
Data selection
scoring criteria
model scalability
low computational cost
deep learning
Data selection
scoring criteria
model scalability
low computational cost
deep learning

期刊

IEEE Transactions on Knowledge and Data Engineering 封面图
IEEE Transactions on Knowledge and Data Engineering
IF:
10.4
论文数:
6.8K
被引数:
3.2W

机构

B
beijing university of posts & telecommunications
学者数:
1.4W
论文数: 1.2W
被引数: 9
U
University of Edinburgh
学者数:
5.2W
论文数: 4.6W
被引数: 71
引用论文

引用论文

Training Data Subset Search With Ensemble Active Learning
err2022-09-01
err8
errOAAI
errChitta, Kashyap; Alvarez, Jose M.; Haussmann, Elmar; Farabet, Clement
err分享
err收藏
err
IF0
err
err0
PREAI
err
err分享
err收藏
err
IF0
err
err0
PREAI
err
err分享
err收藏
学者 查看更多内容