返回
Mango: Exploratory Data Analysis for Large-Scale Sequencing Datasets
DOI:10.1016/j.cels.2019.11.002.png)
摘要
En 中文
The decreasing cost of DNA sequencing over the past decade has led to an explosion of sequencing datasets, leaving us with petabytes of data to analyze. However, current sequencing visualization tools are designed to run on single machines, which limits their scalability and interactivity on modern genomic datasets. Here, we leverage the scalability of Apache Spark to provide Mango, consisting of a Jupyter notebook and genome browser, which removes scalability and interactivity constraints by leveraging multi-node compute clusters to allow interactive analysis over terabytes of sequencing data. We demonstrate scalability of the Mango tools by performing quality control analyses on 10 terabytes of 100 high-coverage sequencing samples from the Simons Genome Diversity Project, enabling capability for interactive genomic exploration of multi-sample datasets that surpass the computational limitations of single-node visualization tools. Mango is freely available for download with full documentation at https://bdg-mango.readthedocs.io/en/latest/.
Keyword:
GENOME BROWSER
HADOOP
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
7.7
论文数:
1.4K
被引数:
1.0W
机构
引用论文
Bioreductive deposition of palladium (0) nanoparticles onShewanella oneidensiswith catalytic activity towards reductive dechlorination of polychlorinated biphenyls钯 (0) 纳米颗粒在 Shewanella oneidensis 上的生物还原沉积,对多氯联苯的还原脱氯具有催化活性
Capacity optimization of hybrid renewable energy system considering part-load ratio and resource endowment考虑部分负荷率和资源禀赋的混合可再生能源系统容量优化

