arrow
返回

Proteogenomic Database Construction Driven from Large Scale RNA-seq Data

delete2013-07-17
delete106
delete
OA
AI
S
Sunghee Woo
S
Seong Won
G
Gennifer E. Merrihew
Y
Yupeng He
N
Natalie Castellana
M
Michael J. MacCoss
V
Vineet Bafna *
DOI:10.1021/pr400294cdelete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
The advent of inexpensive RNA-seq technologies and other deep sequencing technologies for RNA has the promise to radically improve genomic annotation, providing information on transcribed regions and splicing events in a variety of cellular conditions. Using MS-based proteogenomics, many of these events can be confirmed directly at the protein level. However, the integration of large amounts of redundant RNA-seq data and mass spectrometry data poses a challenging problem. Our paper addresses this by construction of a compact database that contains all useful information expressed in RNA-seq reads. Applying our method to cumulative C. elegans data reduced 496.2 GB of aligned RNA-seq SAM files to 410 MB of splice graph database written in FASTA format. This corresponds to 1000x compression of data size, without loss of sensitivity. We performed a proteogenomics study using the custom data set, using a completely automated pipeline, and identified a total of 4044 novel events, including 215 novel genes, 808 novel exons, 12 alternative splicings, 618 gene-boundary corrections, 245 exon-boundary changes, 938 frame shifts, 1166 reverse strands, and 42 translated UTRs. Our results highlight the usefulness of transcript + proteomic integration for improved genome annotations.
Keyword:
proteogenomics
C. elegans
RNA-seq
MS/MS database

期刊

Journal of Proteome Research 封面图
Journal of Proteome Research
IF:
3.6
论文数:
9.3K
被引数:
2.3W

机构

University of California System 封面图
University of California System
学者数:
37.5W
论文数: 33.7W
被引数: 6.6K
U
University of California San Diego
学者数:
4.6W
论文数: 3.5W
被引数: 924