arrow
Return

TIGER: technical variation elimination for metabolomics data using ensemble learning architecture

delete2022-01-03
delete12
delete
OA
AI
S
Siyu Han
J
Jialing Huang
F
Francesco Foppiano
C
Cornelia Prehn
J
Jerzy Adamski
K
Karsten Suhre
Y
Ying Li
G
Giuseppe Matullo
F
Freimut Schliess
C
Christian Gieger
A
Annette Peters
R
Rui Wang‐Sattler *
DOI:10.1093/bib/bbab535delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
Large metabolomics datasets inevitably contain unwanted technical variations which can obscure meaningful biological signals and affect how this information is applied to personalized healthcare. Many methods have been developed to handle unwanted variations. However, the underlying assumptions of many existing methods only hold for a few specific scenarios. Some tools remove technical variations with models trained on quality control (QC) samples which may not generalize well on subject samples. Additionally, almost none of the existing methods supports datasets with multiple types of QC samples, which greatly limits their performance and flexibility. To address these issues, a non-parametric method TIGER (Technical variation elImination with ensemble learninG architEctuRe) is developed in this study and released as an R package (https://CRAN.R-project.org/package=TIGERr). TIGER integrates the random forest algorithm into an adaptable ensemble learning architecture. Evaluation results show that TIGER outperforms four popular methods with respect to robustness and reliability on three human cohort datasets constructed with targeted or untargeted metabolomics data. Additionally, a case study aiming to identify age-associated metabolites is performed to illustrate how TIGER can be used for cross-kit adjustment in a longitudinal analysis with experimental data of three time-points generated by different analytical kits. A dynamic website is developed to help evaluate the performance of TIGER and examine the patterns revealed in our longitudinal analysis (https://han-siyu.github.io/TIGER_web/). Overall, TIGER is expected to be a powerful tool for metabolomics data analysis.
Keywords:
metabolomics
machine learning
ensemble learning
predictive modelling
longitudinal analysis
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

Briefings in Bioinformatics cover
Briefings in Bioinformatics
IF:
7.7
Papers:
5.6K
Citations:
2.7W

Organization

U
University of Turin
Scholars:
3.7W
Papers: 2.8W
Citations: 3.2W
U
University of Munich
Scholars:
5.7W
Papers: 4.2W
Citations: 68
U
University of Ljubljana
Scholars:
1.5W
Papers: 1.3W
Citations: 1.7W
H
Helmholtz Association
Scholars:
13.2W
Papers: 10.7W
Citations: 145
W
Weill Cornell Medicine
Scholars:
2.1W
Papers: 1.5W
Citations: 1.4W
T
Technical University of Munich
Scholars:
5.2W
Papers: 3.9W
Citations: 6.2W
C
Cornell University
Scholars:
6.3W
Papers: 5.4W
Citations: 10.9W
J
Jilin University
Scholars:
8.7W
Papers: 5.5W
Citations: 8.9K
researcher View more organizations