返回
Statistical Binning for Barcoded Reads Improves Downstream Analyses
DOI:10.1016/j.cels.2018.07.005.png)
摘要
En 中文
Sequencing technologies are capturing longer-range genomic information at lower error rates, enabling alignment to genomic regions that are inaccessible with short reads. However, many methods are unable to align reads to much of the genome, recognized as important in disease, and thus report erroneous results in downstream analyses. We introduce EMA, a novel two-tiered statistical binning model for bar-coded read alignment, that first probabilistically maps reads to potentially multiple read clouds'' and then within clouds by newly exploiting the non-uniform read densities characteristic of barcoded read sequencing. EMA substantially improves downstream accuracy over existing methods, including phasing and genotyping on 10x data, with fewer false variant calls in nearly half the time. EMA effectively resolves particularly challenging alignments in genomic regions that contain nearby homologous elements, uncovering variants in the pharmacogenomically important CYP2D region, and clinically important genes C4 (schizophrenia) and AMY1A (obesity), which go undetected by existing methods. Our work provides a framework for future generation sequencing.
Keyword:
DNA-SEQUENCING DATA
HUMAN GENOME
GENERATION
TECHNOLOGIES
FRAMEWORK
ALIGNMENT
VARIANT
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
7.7
论文数:
1.4K
被引数:
1.0W
机构
引用论文
Haplotyping germline and cancer genomes with high-throughput linked-read sequencing
NATURE BIOTECHNOLOGY
IF41.7
The Genome Analysis Toolkit: A MapReduce framework for analyzing next-generation DNA sequencing data基因组分析工具包: 用于分析下一代DNA测序数据的MapReduce框架
GENOME RESEARCH
IF5.5
Are whole-exome and whole-genome sequencing approaches cost-effective? A systematic review of the literature全外显子组和全基因组测序方法是否具有成本效益?对文献的系统回顾
GENETICS IN MEDICINE
IF6.2

