arrow
Return

A data replication algorithm for groups of files in data grids

delete2018-03-01
delete9
PRE
AI
L
Leila Azari *
A
Amir Masoud Rahmani
H
H.A. Daniel
N
Nooruldeen Nasih Qader
DOI:10.1016/j.jpdc.2017.10.008delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Data grid is emerging as the main part of the infrastructure for large-scale data intensive applications such as high energy physics and bioinformatics. The deployment of such infrastructures has allowed users of a grid site to gain access to a large amount of distributed data. Data replication is a key issue in a data grid and could be applied intelligently because it reduces data access time and bandwidth consumption for each grid site. In this paper, we introduce a new dynamic data replication algorithm named Popular Groups of Files Replication (PGFR). Our proposed algorithm is based on an assumption: users in a Virtual Organization have similar interests in groups of files. Based on this assumption, and file access history, PGFR builds a connectivity graph to recognize a group of dependent files in each grid site and replicates the most Popular Groups of Files to each grid site, thus increasing the local availability. We used OptorSim simulator to evaluate the efficiency of PGFR algorithm. The simulation results show that PGFR achieves better performance compared to the existing algorithm; PGFR minimized the mean job execution time, bandwidth consumption, and avoiding unnecessary replication. (C) 2017 Elsevier Inc. All rights reserved.
Keywords:
Dynamic data replication
Data grids
Group replication
Connectivity graph

Journal

Journal of Parallel and Distributed Computing cover
Journal of Parallel and Distributed Computing
IF:
4
Papers:
3.8K
Citations:
4.8K

Organization

I
Islamic Azad University
Scholars:
4.0W
Papers: 3.3W
Citations: 9.8K
U
universidade do algarve
Scholars:
3.9K
Papers: 3.4K
Citations: 6