返回
Reference-free phylogeny from sequencing data
DOI:10.1186/s13040-023-00329-x.png)
摘要
En 中文
MotivationClustering of genetic sequences is one of the key parts of bioinformatics analyses. Resulting phylogenetic trees are beneficial for solving many research questions, including tracing the history of species, studying migration in the past, or tracing a source of a virus outbreak. At the same time, biologists provide more data in the raw form of reads or only on contig-level assembly. Therefore, tools that are able to process those data without supervision need to be developed.ResultsIn this paper, we present a tool for reference-free phylogeny capable of handling data where no mature-level assembly is available. The tool allows distance calculation for raw reads, contigs, and the combination of the latter. The tool provides an estimation of the Levenshtein distance between the sequences, which in turn estimates the number of mutations between the organisms. Compared to the previous research, the novelty of the method lies in a newly proposed combination of the read and contig measures, a new method for read-contig mapping, and an efficient embedding of contigs.
Keyword:
Sequence similarity
Phylogeny
Levenshtein distance
Reads
Contigs
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
6.1
论文数:
700
被引数:
1.5K
机构
引用论文
Genetic Polymorphism of 11 Allozyme Loci in Populations of Wall Lizards (Podarcis sp.) from the Iberian Peninsula and North Africa来自伊比利亚半岛和北非的壁蜥蜴 (Podarcis sp。) 种群中11个同工酶基因座的遗传多态性

