arrow
Return

Distributed Data Deduplication for Big Data: A Survey

delete2025-09-09
delete0
PRE
AI
Y
Yinjin Fu
苏骏 (Jun Su)
J
J.Q. Ning
Y
Yutong Lu
N
Nong Xiao
DOI:10.1145/3735508delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
To address the throughput and capacity limitations of single-node centralized deduplication, distributed data deduplication has become a popular technology in big data management to save more storage space, enhance I/O performance, and improve system scalability. It includes inter-node data assignment from clients to multiple deduplication nodes by a data routing scheme, and independent intra-node redundancy suppression in individual nodes. In this article, we first describe the background of big data deduplication. Then we summarize and classify the state-of-the-art in the key techniques of distributed data deduplication, including data partitioning, chunk fingerprinting, data routing, index lookup, data restoring, garbage collection, the security and reliability of deduplicated data. These help identify and understand the system implementation of the existing distributed data deduplication methods. Moreover, we present some representative industrial products that have successfully applied distributed data deduplication technologies. Finally, we discuss the main challenges and industry trend of distributed data deduplication, and outline the open problems and its future research directions.
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

ACM Computing Surveys cover
ACM Computing Surveys
IF:
28
Papers:
2.4K
Citations:
3.5W

Organization

No organization information available