返回
High-throughput DNA sequence data compression
DOI:10.1093/bib/bbt087.png)
摘要
En 中文
The exponential growth of high-throughput DNA sequence data has posed great challenges to genomic data storage, retrieval and transmission. Compression is a critical tool to address these challenges, where many methods have been developed to reduce the storage size of the genomes and sequencing data (reads, quality scores and metadata). However, genomic data are being generated faster than they could be meaningfully analyzed, leaving a large scope for developing novel compression algorithms that could directly facilitate data analysis beyond data transfer and storage. In this article, we categorize and provide a comprehensive review of the existing compression methods specialized for genomic data and present experimental results on compression ratio, memory usage, time for compression and decompression. We further present the remaining challenges and potential directions for future research.
Keyword:
next-generation sequencing
compression
reference-based compression
reference-free compression
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
7.7
论文数:
5.8K
被引数:
2.7W
机构
引用论文
Reliability and construct validity of the Automated Neuropsychological Assessment Metrics (ANAM) mood scale自动神经心理学评估指标 (ANAM) 情绪量表的信度和结构效度
Compression of next-generation sequencing reads aided by highly efficient de novo assembly
NUCLEIC ACIDS RESEARCH
IF13.1

