arrow
Return

File aware distributed deduplication system in cloud environment

delete2025-11-15
delete0
PRE
AI
A
Amdewar Godavari *
C
Chapram Sudhakar
T
T. Ramesh
DOI:10.1007/s11227-025-08058-0delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
In the cloud environment, number of clients are generating huge amounts of duplicate content, in the form of different types of files. Data redundancy is more prevalent among the same type of files, and it is negligible across different types of files. Applying the same deduplication technique irrespective of the file type results in wastage of resources, with insignificant duplicates elimination. If the same type of files is routed to the same data server, with deduplication, more duplicate content can be eliminated, but raises the storage imbalance problem. If files are distributed among data servers, ignoring the file type, better storage balance can be achieved, but the eliminated duplicate content decreases. In order to address these challenges, a distributed deduplication system is proposed. The system categorizes files into three groups—high, low, and unpredictable duplicate files—based on their redundancy levels. A group of data servers is allocated for each category of files. Preprocessing at source and resource-intensive duplicate content identification and elimination at the data servers reduces communication overhead. Data servers apply similarity-based segment-level deduplication for high and unpredictable duplicate files, and file-level deduplication is applied for low duplicate files. This system gives a high degree of load balance while achieving comparable space saving.
Keywords:
Distributed deduplication
Disk bottleneck
Load balancing
Cloud computing
Cloud storage

Journal

T
The Journal of Supercomputing
IF:
0
Papers:
647
Citations:
0

Organization

D
Department of Computer Science
Scholars:
1.7K
Papers: 998
Citations: 8
Cited Papers

Cited Papers

A study of practical deduplication
err2012-02-02
err0
errOAAI
errDutch T. Meyer; William J. Bolosky
errShare
errSave
Boafft: Distributed Deduplication for Big Data Storage in the Cloud
err2020-10-01
err30
errOAAI
errLuo, Shengmei; Zhang, Guangyan; Wu, Chengwen; Khan, Samee U.; Li, Keqin
errShare
errSave
errShare
errSave
D3: A Dynamic Dual-Phase Deduplication Framework for Distributed Primary Storage
err2018-02-01
err8
PREAI
errYin, Jianwei; Tang, Yan; Deng, Shuiguang; Li, Ying; Zomaya, Albert Y.
errShare
errSave
FASR: An Efficient Feature-Aware Deduplication Method in Distributed Storage Systems
err2022-01-01
err5
errOAAI
errYao, Wenbin; Hao, Mengyao; Hou, Yingying; Li, Xiaoyong
errShare
errSave
no more