arrow
Return

Optimal compressed representation of high throughput sequence data via light assembly

delete2018-02-08
delete15
delete
OA
AI
A
Antonio A. Ginart
J
Joseph Hui
K
Kaiyuan Zhu
I
Ibrahim Numanagić
T
Thomas A. Courtade
S
S. Cenk Şahinalp *
D
David Tse
DOI:10.1038/s41467-017-02480-6delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
The most effective genomic data compression methods either assemble reads into contigs, or replace them with their alignment positions on a reference genome. Such methods require significant computational resources, but faster alternatives that avoid using explicit or de novo-constructed references fail to match their performance. Here, we introduce a new reference-free compressed representation for genomic data based on light de novo assembly of reads, where each read is represented as a node in a (compact) trie. We show how to efficiently build such tries to compactly represent reads and demonstrate that among all methods using this representation (including all de novo assembly based methods), our method achieves the shortest possible output. We also provide an lower bound on the compression rate achievable on uniformly sampled genomic read data, which is approximated by our method well. Our method significantly improves the compression performance of alternatives without compromising speed.
Keywords:
ALGORITHM
SETS
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

Nature Communications cover
Nature Communications
IF:
15.7
Papers:
9.2W
Citations:
91.2W

Organization

I
indiana university system
Scholars:
4.0W
Papers: 3.5W
Citations: 38
S
Stanford University
Scholars:
9.6W
Papers: 8.2W
Citations: 17.0W
I
Indiana University Bloomington
Scholars:
1.9W
Papers: 1.5W
Citations: 2.8W
researcher View more organizations