arrow
Return

Data Preparation for Large Language Models

delete2026-03-01
delete0
PRE
AI
H
Hao Liang
Z
Zhen Hao Wong
L
Liu, Rui-Tong
Y
Yu-Han Wang
M
Mei-Yi Qiang
Z
Zheng-Yang Zhao
C
Cheng-Yu Shen
C
Cong-hui He
W
Wen-Tao Zhang *
B
Bin Cui *
DOI:10.1007/s11390-026-5948-8delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Large language models (LLMs) have demonstrated remarkable generalization capabilities across diverse domains, largely attributed to the availability of massive amounts of high-quality training data. Recently, the development paradigm of LLMs has been shifting from a model-centric to a data-centric perspective. In this paper, we provide a comprehensive survey of data preparation algorithms and workflows for LLMs, categorized into three stages: pre-training, continual pre-training, and post-training. We further summarize widely used datasets along with their associated data preparation method, offering a practical reference for researchers who may lack extensive experience in the field of data preparation. Finally, we outline potential directions for future work, highlighting open challenges and opportunities in advancing data preparation for LLMs.
Keywords:
data-centric artificial intelligence (AI)
data management
large language model (LLM)

Journal

J
Journal of Computer Science and Technology
IF:
1.3
Papers:
61
Citations:
1.5K

Organization

S
Shanghai Artificial Intelligence Laboratory
Scholars:
457
Papers: 256
Citations: 765
P
peking university
Scholars:
11.7W
Papers: 8.7W
Citations: 146