arrow
Return

Utilizing sequence intrinsic composition to classify protein-coding and long non-coding transcripts

delete2013-07-27
delete1.5K
delete
OA
AI
孙亮 (Liang Sun)
H
Haitao Luo
D
Dechao Bu
G
Guoguang Zhao
K
Kuntao Yu
C
Changhai Zhang
Y
Yuanning Liu
陈润生 (Runsheng Chen)
赵屹 (Yi Zhao) *
DOI:10.1093/nar/gkt646delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
It is a challenge to classify protein-coding or non-coding transcripts, especially those re-constructed from high-throughput sequencing data of poorly annotated species. This study developed and evaluated a powerful signature tool, Coding-Non-Coding Index (CNCI), by profiling adjoining nucleotide triplets to effectively distinguish protein-coding and non-coding sequences independent of known annotations. CNCI is effective for classifying incomplete transcripts and sense-antisense pairs. The implementation of CNCI offered highly accurate classification of transcripts assembled from whole-transcriptome sequencing data in a cross-species manner, that demonstrated gene evolutionary divergence between vertebrates, and invertebrates, or between plants, and provided a long non-coding RNA catalog of orangutan. CNCI software is available at http://www.bioinfo.org/software/cnci.
Keywords:
INTEGRATIVE ANNOTATION
GENE
RNAS
PREDICTION
REVEALS
EVOLUTION
TOPHAT
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

Nucleic Acids Research cover
Nucleic Acids Research
IF:
13.1
Papers:
3.6W
Citations:
29.0W

Organization

I
institute of computing technology, cas
Scholars:
1.0K
Papers: 877
Citations: 1
J
Jilin University
Scholars:
8.7W
Papers: 5.5W
Citations: 8.9K
C
chinese academy of sciences
Scholars:
56.5W
Papers: 44.9W
Citations: 704
researcher View more organizations