1
Return

Detecting Training Data For Large Language Models: A Survey

delete2026-05-23
delete0
PRE
AI
C
Chen Yang *
J
Junyi Li
S
Shulin LAN
Y
Yingchao Wang
H
Hongyang Du
X
Xingshan Yao
D
Dusit Niyato
L
Liehuang Zhu
DOI:10.1145/3779430delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
As large language models (LLMs) continue to evolve, the scope and diversity of data used for training are expanding significantly. However, the training dataset of LLMs may inevitably contain sensitive information such as personal data or copyrighted material, leading to privacy leakage or copyright infringement risks if the model generates highly similar or identical text to these sources. This has drawn attention to the issue of detecting whether the text data is used for LLM training. To date, research on detecting training data usage in artificial intelligence (AI) models has mainly focused on traditional machine learning (ML) models. However, studies on LLMs remain relatively immature. The lack of understanding of research progress in this area has hindered the development of more effective detection methods. Therefore, this article aims to address this gap by conducting the analysis of detecting training data for LLM. Specifically, we analyze the available LLM's information to the detector, the main detection methods, determination metrics, and discuss the technical challenges and potential directions for future research in this field.
Keywords:
Large language models
detecting training data

Journal

ACM Computing Surveys cover
ACM Computing Surveys
IF:
28
Papers:
2.4K
Citations:
3.5W

Organization

U
university of chinese academy of sciences, cas
Scholars:
4.1W
Papers: 3.8W
Citations: 74
U
university of hong kong
Scholars:
3.0K
Papers: 1.4K
Citations: 0
B
beijing institute of technology
Scholars:
5.3W
Papers: 3.9W
Citations: 63
C
chinese academy of sciences
Scholars:
54.9W
Papers: 44.5W
Citations: 703
Cited Papers

Cited Papers

Citing Papers

Citing Papers