Return
Distributed Data Deduplication for Big Data: A Survey
DOI:10.1145/3735508.png)
Abstract
En 中文
To address the throughput and capacity limitations of single-node centralized deduplication, distributed data deduplication has become a popular technology in big data management to save more storage space, enhance I/O performance, and improve system scalability. It includes inter-node data assignment from clients to multiple deduplication nodes by a data routing scheme, and independent intra-node redundancy suppression in individual nodes. In this article, we first describe the background of big data deduplication. Then we summarize and classify the state-of-the-art in the key techniques of distributed data deduplication, including data partitioning, chunk fingerprinting, data routing, index lookup, data restoring, garbage collection, the security and reliability of deduplicated data. These help identify and understand the system implementation of the existing distributed data deduplication methods. Moreover, we present some representative industrial products that have successfully applied distributed data deduplication technologies. Finally, we discuss the main challenges and industry trend of distributed data deduplication, and outline the open problems and its future research directions.
AI Summary
Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.
Journal
IF:
28
Papers:
2.4K
Citations:
3.5W
Organization
No organization information available

