arrow
Return

Cross-Study Replicability in Cluster Analysis

delete2023-05-01
delete2
delete
OA
AI
L
Lorenzo Masoero *
E
Emma G. Thomas
G
Giovanni Parmigiani
S
Svitlana Tyekucheva
L
Lorenzo Trippa
DOI:10.1214/22-STS871delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
In cancer research, clustering techniques are widely used for ex-ploratory analyses, playing a critical role in the identification of novel cancer subtypes and patient management. As data collected by multiple research groups grows, it is increasingly feasible to investigate the replicability of clustering procedures, that is, their ability to consistently recover biologi-cally meaningful clusters across several data sets. In this paper, we review methods for replicability of clustering analyses, and discuss a novel frame-work for evaluating cross-study clustering replicability, useful when two or more studies are available. Our approach can be applied to any clustering al-gorithm and can employ different measures of similarity between partitions to quantify replicability, globally (i.e., for the whole sample) as well as lo-cally (i.e., for individual clusters). Using experiments on synthetic and real gene expression data, we illustrate the usefulness of our procedure to evalu-ate if the same clusters are identified consistently across a collection of data sets.
Keywords:
Clustering
replicability
multiple studies

Journal

Statistical Science cover
Statistical Science
IF:
3.4
Papers:
1.0K
Citations:
8.7K

Organization

R
rand health
Scholars:
171
Papers: 124
Citations: 0
RAND Corporation cover
RAND Corporation
Scholars:
2.8K
Papers: 3.4K
Citations: 3.0K
A
amazon.com
Scholars:
698
Papers: 505
Citations: 8
researcher View more organizations