Return
Statistical analysis of correlated expression data from high throughput experiments
DOI:10.1093/genetics/iyaf060.png)
Abstract
En 中文
Data obtained from high throughput experiments often exhibit complex dependencies among features. These dependencies arise from various sources, including genetic correlation, batch effects, technical replicates, and shared biological pathways. Ignoring these dependencies can lead to inflated false discovery rate (FDR), reduced statistical power, and biased biological interpretations. Properly accounting for these dependencies is crucial for accurate detection of biological signals. We propose a new method called Analysis of Correlated Expressions (ACE) to compare the mean expression of features between two groups. ACE is based on a factor analytic model that accounts for dependence among features and also incorporates heterogeneity of variances between groups, a common feature of high throughput data. Furthermore, ACE does not require the data to be normally distributed. It is scalable and free of any unknown tuning parameters. Extensive simulation studies indicate that it is more powerful than many existing methods while controlling the FDR. Application of ACE to a microRNA dataset, a neuroblastoma gene expression dataset, and a Huntington's disease dataset resulted in some novel findings that were missed by existing methods.
Keywords:
dependence
factor model
false discovery rate
gene expression
high throughput experiment
multiple testing

