arrow
返回

A fast and efficient algorithm for DNA sequence similarity identification

delete2022-08-23
delete1
delete
OA
AI
M
Machbah Uddin
M
Mohammad Khairul Islam *
M
Md. Rakib Hassan
F
Farah Jahan
J
Joong-Hwan Baek
DOI:10.1007/s40747-022-00846-ydelete
delete原文链接
delete分享
delete收藏
查看原文
摘要

摘要

En 中文
DNA sequence similarity analysis is necessary for enormous purposes including genome analysis, extracting biological information, finding the evolutionary relationship of species. There are two types of sequence analysis which are alignment-based (AB) and alignment-free (AF). AB is effective for small homologous sequences but becomes N P -hard problem for long sequences. However, AF algorithms can solve the major limitations of AB. But most of the existing AF methods show high time complexity and memory consumption, less precision, and less performance on benchmark datasets. To minimize these limitations, we develop an AF algorithm using a 2D k -mer count matrix inspired by the CGR approach. Then we shrink the matrix by analyzing the neighbors and then measure similarities using the best combinations of pairwise distance (PD) and phylogenetic tree methods. We also dynamically choose the value of k for k - mer. We develop an efficient system for finding the positions of k - mer in the count matrix. We apply our system in six different datasets. We achieve the top rank for two benchmark datasets from AFproject, 100% accuracy for two datasets (16 S Ribosomal, 18 Eutherian), and achieve a milestone for time complexity and memory consumption in comparison to the existing study datasets (HEV, HIV-1). Therefore, the comparative results of the benchmark datasets and existing studies demonstrate that our method is highly effective, efficient, and accurate. Thus, our method can be used with the top level of authenticity for DNA sequence similarity measurement.
Keyword:
DNA sequence similarity
Dynamic k - k - mer count matrix
Matrix shrinking
AFproject
Benchmark dataset
Bioinformatics engineering
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

Complex and Intelligent Systems 封面图
Complex and Intelligent Systems
IF:
4.6
论文数:
2.1K
被引数:
6.6K

机构

B
bangladesh agricultural university (bau)
学者数:
2.4K
论文数: 1.3K
被引数: 2
K
Korea Aerospace University
学者数:
1.1K
论文数: 1.1K
被引数: 513
U
University of Chittagong
学者数:
1.6K
论文数: 1.0K
被引数: 1.3K
学者 查看更多机构
引用论文

引用论文

Ambulatory blood pressure and subclinical cardiovascular disease in patients with juvenile-onset systemic lupus erythematosus
err2012-10-10
err0
PREAI
errNur Canpolat; Ozgur Kasapcopur; Salim Caliskan; Selman Gokalp; Meltem Bor; Mehmet Tasdemir; Lale Sever; Nil Arisoy
err分享
err收藏
Bail-ins and Bail-outs: Incentives, Connectivity, and Systemic Stability
err
IF0
err2017-08-01
err0
errOAAI
errBenjamin Bernard; Agostino Capponi; Joseph Stiglitz
err分享
err收藏
Erste Ergebnisse zu Reliabilität und Validität der OPD-2 Strukturachse
err2009-02-01
err0
PREAI
errCord Benecke; Andrea Koschier; Doris Peham; Astrid Bock; Reiner W. Dahlbender; Wilfried Biebl; Stephan Doering
err分享
err收藏
Alignment-free sequence comparison: benefits, applications, and tools
err2017-10-03
err336
errOAAI
errZielezinski, Andrzej; Vinga, Susana; Almeida, Jonas; Karlowski, Wojciech M.
err分享
err收藏
CAFE: aCcelerated Alignment-FrEe sequence analysis
err2017-05-03
err51
errOAAI
errLu, Yang Young; Tang, Kujin; Ren, Jie; Fuhrman, Jed A.; Waterman, Michael S.; Sun, Fengzhu
err分享
err收藏
学者 查看更多内容