Return
DigestedProteinDB: A Compact and Scalable Key-Value Database for In Silico Peptide Digestion and Mass-Based Search
T
J
J
K
A
A
DOI:10.1021/acs.jproteome.6c00152.png)
Abstract
En 中文
Efficient comparison of experimental peptide masses with theoretical values is a central step in mass spectrometry (MS)–based proteomics, microbial biotyping, and MS imaging. Such analyses increasingly require fast and scalable access to large collections of in silico–digested peptides derived from large-scale and continuously evolving protein-sequence databases. Here, we present DigestedProteinDB, a compact and high-performance key–value database of peptides generated by enzymatic in silico digestion of UniProtKB/Swiss-Prot and TrEMBL sequences. Implemented using RocksDB, the system incorporates multiple optimization layers, such as peptide mass discretization and multistage storage compression, to minimize disk footprint and accelerate mass-range queries. In benchmark tests using 252 million UniProtKB protein sequences (5.9 billion peptides, trypsin; 6–50 aa; two missed cleavages), DigestedProteinDB required approximately 250 GB of disk space and operated within 16 GB of system RAM. Database construction required ∼2 days, and end-to-end mass-range queries (±0.1 Da) achieved a median latency of ∼7 ms per query across a batch of 10,000 randomly sampled queries. The resulting database can be used as a standalone local resource or integrated into bioinformatics pipelines for peptide mass fingerprinting (PMF), MS/MS-based protein identification, and microbial biotyping. Due to its modular design, new databases can be generated rapidly for different proteases, taxonomic subsets, or digestion parameters.
Keywords:
Biological databases
Circuits
Compression
Mass spectrometry
Peptides and proteins
in silico digestion
mass spectrometry
peptide database
peptide mass fingerprinting
UniProt
RocksDB
protein identification
bioinformatics
Journal
IF:
3.6
Papers:
9.3K
Citations:
2.3W

