返回
Efficient Motif Discovery in Protein Sequences Using a Branch and Bound Algorithm
DOI:10.1109/JBHI.2024.3355964.png)
摘要
En 中文
Identifying motifs within sets of protein sequences constitutes a pivotal challenge in proteomics, imparting insights into protein evolution, function prediction, and structural attributes. Motifs hold the potential to unveil crucial protein aspects like transcription factor binding sites and protein-protein interaction regions. However, prevailing techniques for identifying motif sequences in extensive protein collections often entail significant time investments. Furthermore, ensuring the accuracy of obtained results remains a persistent motif discovery challenge. This paper introduces an innovative approach-a branch and bound algorithm-for exact motif identification across diverse lengths. This algorithm exhibits superior performance in terms of reduced runtime and enhanced result accuracy, as compared to existing methods. To achieve this objective, the study constructs a comprehensive tree structure encompassing potential motif evolution pathways. Subsequently, the tree is pruned based on motif length and targeted similarity thresholds. The proposed algorithm efficiently identifies all potential motif subsequences, characterized by maximal similarity, within expansive protein sequence datasets. Experimental findings affirm the algorithm's efficacy, highlighting its superior performance in terms of runtime, motif count, and accuracy, in comparison to prevalent practical techniques.
Keyword:
Proteins
Protein engineering
Biology
Bioinformatics
Search methods
Amino acids
Runtime
branch and bound approach
motif discovery
protein sequences
tree-based algorithm
期刊
IF:
6.8
论文数:
4.6K
被引数:
2.0W
机构
引用论文
A graph-based motif detection algorithm models complex nucleotide dependencies in transcription factor binding sites
NUCLEIC ACIDS RESEARCH
IF13.1
Computer-aided design of RNA-targeted small molecules: A growing need in drug discoveryRNA靶向小分子的计算机辅助设计: 药物发现中日益增长的需求
CHEM
IF19.6
A unified catalog of 204,938 reference genomes from the human gut microbiome来自人类肠道微生物组的204,938参考基因组的统一目录
NATURE BIOTECHNOLOGY
IF41.7

