返回
Accelerating read mapping with FastHASH
DOI:10.1186/1471-2164-14-S1-S13.png)
摘要
En 中文
With the introduction of next-generation sequencing (NGS) technologies, we are facing an exponential increase in the amount of genomic sequence data. The success of all medical and genetic applications of next-generation sequencing critically depends on the existence of computational techniques that can process and analyze the enormous amount of sequence data quickly and accurately. Unfortunately, the current read mapping algorithms have difficulties in coping with the massive amounts of data generated by NGS. We propose a new algorithm, FastHASH, which drastically improves the performance of the seed-and-extend type hash table based read mapping algorithms, while maintaining the high sensitivity and comprehensiveness of such methods. FastHASH is a generic algorithm compatible with all seed-and-extend class read mapping algorithms. It introduces two main techniques, namely Adjacency Filtering, and Cheap K-mer Selection. We implemented FastHASH and merged it into the codebase of the popular read mapping program, mrFAST. Depending on the edit distance cutoffs, we observed up to 19-fold speedup while still maintaining 100% sensitivity and high comprehensiveness.
Keyword:
SEGMENTAL DUPLICATIONS
COPY NUMBER
GENOME
EVOLUTION
ALIGNMENT
SEQUENCES
SEARCH
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
3.7
论文数:
1.9W
被引数:
5.2W
机构
引用论文
A large and complex structural polymorphism at 16p12.1 underlies microdeletion disease risk16 p12.1处的大而复杂的结构多态性是微缺失疾病风险的基础
NATURE GENETICS
IF31.8
Gene expression analysis by massively parallel signature sequencing (MPSS) on microbead arrays
NATURE BIOTECHNOLOGY
IF41.7
Characterization of polymer matrix and low melting point solder for anisotropic conductive film各向异性导电膜用聚合物基体及低熔点焊料的表征
Segmental duplications: Organization and impact within the current Human Genome Project assembly
GENOME RESEARCH
IF5.5

