Return
Managing Big Interval Data with CINTIA: The Checkpoint INTerval Array
DOI:10.1109/TBDATA.2017.2691719.png)
Abstract
En 中文
Intervals have become prominent in data management as they are the main data structure to represent a number of key data types such as temporal or genomic data. Yet, there exists no solution to compactly store and efficiently query big interval data. In this paper we introduce CINTIA-the Checkpoint INTerval Index Array-an efficient data structure to store and query interval data, which achieves high memory locality and outperforms state-of-the art solutions. We also propose a low-latency, Big Data system that implements CINTIA on top of a popular distributed file system and efficiently manages large interval data on clusters of commodity machines. Our system can easily be scaled-out and was designed to accommodate large delays between the various components of a distributed infrastructure. We experimentally evaluate the performance of our approach on several datasets and show that it outperforms current solutions by several orders of magnitude in distributed settings.
Keywords:
Indexes
Arrays
Complexity theory
Big Data
Bioinformatics
Genomics
Interval data
low-latency
scalability
distributed data management
AI Summary
Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.
Journal
I
IF:
5.7
Papers:
860
Citations:
3.0K

