arrow
Return

Using Markov model to improve word normalization algorithm for biological sequence comparison

delete2011-04-20
delete3
PRE
AI
Q
Qi Dai *
X
Xiaoqing Liu
Y
Yuhua Yao
F
Fukun Zhao
DOI:10.1007/s00726-011-0906-2delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
There are two crucial problems with statistical measures for sequence comparison: overlapping structures and background information of words in biological sequences. Word normalization in improved composition vector method took into account these problems and achieved better performance in evolutionary analysis. The word normalization is desirable, but not sufficient, because it assumes that the four bases A, C, T, and G occur randomly with equal chance. This paper proposed an improved word normalization which uses Markov model to estimate exact k-word distribution according to observed biological sequence and thus has the ability to adjust the background information of the k-word frequencies in biological sequences. The improved word normalization was tested with three experiments and compared with the existing word normalization. The experiment results confirm that the improved word normalization using Markov model to estimate the exact k-word distribution in biological sequences is more efficient.
Keywords:
Markov model
Word normalization
Sequence comparison
Classification
Phylogenetic analysis
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

Amino Acids cover
Amino Acids
IF:
2.4
Papers:
4.7K
Citations:
1.0W

Organization

H
Hangzhou Dianzi University
Scholars:
1.3W
Papers: 9.6K
Citations: 7.5K
Z
Zhejiang Sci-Tech University
Scholars:
1.7W
Papers: 1.0W
Citations: 1.3W