arrow
Return

Optimizing differential expression analysis for proteomics data via high-performing rules and ensemble inference

delete2024-05-09
delete2
delete
OA
AI
H
Hui Peng
H
He Wang
W
Weijia Kong
李金燕 (Jinyan Li) *
W
Wilson Wen Bin Goh *
DOI:10.1038/s41467-024-47899-wdelete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
Identification of differentially expressed proteins in a proteomics workflow typically encompasses five key steps: raw data quantification, expression matrix construction, matrix normalization, missing value imputation (MVI), and differential expression analysis. The plethora of options in each step makes it challenging to identify optimal workflows that maximize the identification of differentially expressed proteins. To identify optimal workflows and their common properties, we conduct an extensive study involving 34,576 combinatoric experiments on 24 gold standard spike-in datasets. Applying frequent pattern mining techniques to top-ranked workflows, we uncover high-performing rules that demonstrate optimality has conserved properties. Via machine learning, we confirm optimal workflows are indeed predictable, with average cross-validation F1 scores and Matthew's correlation coefficients surpassing 0.84. We introduce an ensemble inference to integrate results from individual top-performing workflows for expanding differential proteome coverage and resolve inconsistencies. Ensemble inference provides gains in pAUC (up to 4.61%) and G-mean (up to 11.14%) and facilitates effective aggregation of information across varied quantification approaches such as topN, directLFQ, MaxLFQ intensities, and spectral counts. However, further development and evaluation are needed to establish acceptable frameworks for conducting ensemble inference on multiple proteomics workflows. In proteomics, identifying differentially expressed proteins (DEPs) is critical for uncovering biomarkers and drug targets. However, constructing optimal workflows to achieve maximal identification of DEPs is challenging. Here, the authors performed 34,576 combinatorial experiments on 24 gold standard spike-in datasets to discern optimal workflows.
Keywords:
MASS-SPECTROMETRY
PEPTIDE IDENTIFICATION
NORMALIZATION METHODS
R-PACKAGE
BENCHMARKING
IMPUTATION
VARIANCE
SOFTWARE
STANDARD
QUANTIFICATION
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

Nature Communications cover
Nature Communications
IF:
15.7
Papers:
9.3W
Citations:
91.2W

Organization

N
Nanyang Technological University
Scholars:
4.9W
Papers: 4.8W
Citations: 8.1W
C
chinese academy of sciences
Scholars:
56.5W
Papers: 44.9W
Citations: 704