Return
Evaluating statistical models for overdispersed multi-omics data: a multiplex immunofluorescence case study
DOI:10.1093/aje/kwag127.png)
Abstract
En 中文
Multi-omic data analysis poses statistical challenges. We evaluated statistical models for our multiplex immunofluorescence study of T cell subset densities in colorectal cancer. Using 1235 cases, we compared seven models- ordinal logistic regression, Poisson, quasi-Poisson, quadratic negative binomial (NB), linear NB, zero-inflated NB (ZINB) and hurdle NB models- assessing associations with a strong (microsatellite instability, MSI) and a weak (calcium intake) exposure. Simulation studies assessed type I error and power. Effect estimates were generally consistent for the strong exposure (MSI) but varied for the weaker exposure (calcium). Simulations revealed inflated false-positive rates for the Poisson and NB-based models, including quadratic NB, zero-inflated and hurdle, but not for ordinal logistic regression or the linear NB model. The quasi-Poisson model showed modest inflation of low p-values, but the overall p-value distribution remained approximately uniform under the null. Ordinal logistic, linear NB, and quasi-Poisson models achieved the best or near-best power across a range of zero proportions, dispersion levels, and distributions. The ordinal logistic, linear NB, and quasi-Poisson models are useful and robust options for epidemiologic analyses of overdispersed, right-skewed multi-omic data with a nontrivial proportion of zero counts.
Journal
IF:
4.8
Papers:
9.9K
Citations:
3.7W

