Return
Robust Microbial Signature Discovery via Post-Selection Inference for Microbiome Compositions
W
X
H
王
DOI:10.1080/01621459.2026.2671446.png)
Abstract
En 中文
Identifying taxa associated with host phenotypes is crucial for understanding host-microbe interactions and their underlying molecular mechanisms. However, analyzing microbiome data presents unique challenges, as the observed abundances of taxa are high-dimensional, compositional, and subject to both sample-specific and taxon-specific biases. Many existing methods for differential abundance testing struggle to balance false discovery rate control with statistical power. In this article, we propose PoDA, a post-selection inference method for differential abundance analysis to address the limitations of the existing methods. PoDA begins by selecting a subset of taxa likely associated with the phenotype using penalized regression under a mean-shift model. It then leverages the unselected taxa to correct for sample-specific bias and assign p-values to the selected taxa. To ensure valid inference after the selection process, PoDA employs an information-splitting procedure, which is repeated to enhance stability and power. Comprehensive simulation studies demonstrate the superiority of PoDA over existing methods. We further applied PoDA to two microbiome case-control studies of Parkinson’s disease (PD). The method identified a set of candidate microbial signatures associated with PD and showed improved replicability across the two independent datasets. PoDA is implemented in R and available at https://github.com/weihaowang01/PoDA. Supplementary materials for this article are available online, including a standardized description of the materials available for reproducing the work.
Keywords:
Compositional bias
Data splitting
Mean-shift models
Post-selection inference
Summary statistics
Journal
J
IF:
3
Papers:
5.1K
Citations:
4.8W
