arrow
Return

NASCUP: Nucleic Acid Sequence Classification by Universal Probability

delete2021-01-01
delete1
delete
OA
AI
S
Sunyoung Kwon
G
Gyuwan Kim
B
Byunghan Lee
J
Jongsik Chun
S
Sungroh Yoon *
Y
Young-Han Kim *
DOI:10.1109/ACCESS.2021.3127957delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Nucleic acid sequence classification is a fundamental task in the field of bioinformatics. Due to the increasing amount of unlabeled nucleotide sequences, fast and accurate classification of them on a large scale has become crucial. In this work, we developed NASCUP, a new classification method that captures statistical structures of nucleotide sequences by compact context-tree models and universal probability from information theory. A comprehensive experimental study involving nine public databases for functional non-coding RNA, microbial taxonomy and coding/non-coding RNA classification demonstrates the advantages of NASCUP over widely-used alternatives in efficiency, accuracy, and scalability across all datasets considered. NASCUP achieved BLAST-like classification accuracy consistently for several large-scale databases in orders-of-magnitude reduced runtime, and was applied to other bioinformatics tasks such as outlier detection and synthetic sequence generation.
Keywords:
Context modeling
Markov processes
Hidden Markov models
Data models
Maximum likelihood estimation
Probability
Databases
Bioinformatics
context-tree models
information theory
sequence classification
universal probability

Journal

IEEE Access cover
IEEE Access
IF:
3.6
Papers:
9.8W
Citations:
29.4W

Organization

U
University of California Santa Barbara
Scholars:
1.2W
Papers: 9.6K
Citations: 3.6W
P
pusan national university
Scholars:
2.1W
Papers: 1.9W
Citations: 20
University of California System cover
University of California System
Scholars:
37.5W
Papers: 33.7W
Citations: 6.6K
S
seoul national university (snu)
Scholars:
7.2W
Papers: 6.6W
Citations: 86
researcher View more organizations