arrow
Return

Dictionary-driven prokaryotic gene finding

delete2002-06-15
delete25
delete
OA
AI
T
Tetsuo Shibuya
R
Rigoutsos, I
DOI:10.1093/nar/gkf338delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
Gene Identification, also known as gene finding or gene recognition, Is among the important problems of molecular biology that have been receiving Increasing attention with the advent of large scale sequencing projects. Previous strategies for solving this problem can be categorized into essentially two schools of thought: one school employs sequence composition statistics, whereas the other relies on database similarity searches. In this paper, we propose a new gene Identification scheme that combines the best characteristics from each of these two schools. In particular, our method determines gene candidates among the ORFs that can be identified In a given DNA strand through the use of the Bio-Dictionary, a database of patterns that covers essentially all of the currently available sample of the natural protein sequence space. Our approach relies entirely on the use of redundant patterns as the agents on which the presence or absence of genes Is predicated and does not employ any additional evidence, e.g. ribosome-binding site signals. The Bio-Dictionary Gene Finder (BDGF), the algorithm's implementation, is a single computational engine able to handle the gene identification task across distinct archaeal and bacterial genomes. The engine exhibits performance that is characterized by simultaneous very high values of sensitivity and specificity, and a high percentage of correctly predicted start sites. Using a collection of patterns derived from an old (June 2000) release of the Swiss-Prot/TrEMBL database that contained 451 602 proteins and fragments, we demonstrate our method's generality and capabilities through an extensive analysis of 17 complete archaeal and bacterial genomes. Examples of previously unreported genes are also shown and discussed In detail.
Keywords:
PROTEIN-CODING REGIONS
COMPUTATIONAL METHODS
PATTERN DISCOVERY
IDENTIFICATION
DATABASE
RECOGNITION
ALGORITHM
SEQUENCES
ALIGNMENT
STARTS
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

Nucleic Acids Research cover
Nucleic Acids Research
IF:
13.1
Papers:
3.6W
Citations:
29.0W

Organization

No organization information available