arrow
Return

Managing Big Interval Data with CINTIA: The Checkpoint INTerval Array

delete2021-06-01
delete0
PRE
AI
R
Ruslan Mavlyutov *
P
Philippe Cudré-Mauroux
DOI:10.1109/TBDATA.2017.2691719delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Intervals have become prominent in data management as they are the main data structure to represent a number of key data types such as temporal or genomic data. Yet, there exists no solution to compactly store and efficiently query big interval data. In this paper we introduce CINTIA-the Checkpoint INTerval Index Array-an efficient data structure to store and query interval data, which achieves high memory locality and outperforms state-of-the art solutions. We also propose a low-latency, Big Data system that implements CINTIA on top of a popular distributed file system and efficiently manages large interval data on clusters of commodity machines. Our system can easily be scaled-out and was designed to accommodate large delays between the various components of a distributed infrastructure. We experimentally evaluate the performance of our approach on several datasets and show that it outperforms current solutions by several orders of magnitude in distributed settings.
Keywords:
Indexes
Arrays
Complexity theory
Big Data
Bioinformatics
Genomics
Interval data
low-latency
scalability
distributed data management
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

I
IEEE Transactions on Big Data
IF:
5.7
Papers:
860
Citations:
3.0K

Organization

U
University of Fribourg
Scholars:
4.7K
Papers: 4.0K
Citations: 7.5K