arrow
Return

Machine Learning-Based Data Deduplication: Techniques, Challenges, and Future Directions

delete2026-02-01
delete0
PRE
AI
R
Ravneet Kaur *
H
Harcharan Jit Singh
I
Inderveer Chana
DOI:10.1002/cpe.70574delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Data deduplication plays an important role in modern data management as it reduces storage costs and ensures consistency by eliminating redundant records. The traditional data deduplication methods are effective for exact matches but struggle with adaptability and detecting near-exact duplicate records in unstructured or complex data. Machine learning (ML) addresses these limitations by using pattern recognition, feature learning, and statistical modeling to identify subtle similarities between records. This review classifies ML-based deduplication techniques into supervised, unsupervised, semi-supervised, and deep learning methodologies. It also discusses key challenges, including class imbalance, model interpretability, and computational overhead. The paper also explores recent developments in federated learning, real-time deduplication, and multimodal techniques to highlight current trends in these areas. Finally, the paper identifies key open issues and proposes a unified perspective for scalable, real-time deduplication systems that can accommodate diverse data types, structures, and system requirements.
Keywords:
data deduplication
deep learning
machine learning
similarity matching
storage optimization
structured and unstructured data

Journal

C
CONCURRENCY AND COMPUTATION-PRACTICE & EXPERIENCE
IF:
1.5
Papers:
473
Citations:
0

Organization

T
thapar institute of engineering & technology
Scholars:
332
Papers: 168
Citations: 0