Return
STEAM: Spatial Transcriptomics Evaluation Algorithm and Metric for clustering performance
DOI:10.1093/bib/bbaf570.png)
Abstract
En 中文
Spatial transcriptomic technologies allow researchers to explore the diversity and specificity of gene expression within an original tissue structure. Accurately identifying regions that are spatially coherent in both gene expression and physical tissue structures is an emerging topic; however, it is challenging because of the lack of ground truth labels which complicates the validation of clustering consistency and reproducibility. This highlights a need for a computational evaluation framework to rigorously and unbiasedly assess the clustering performance. To address this gap, we propose Spatial Transcriptomics Evaluation Algorithm and Metric (STEAM), a user-friendly computational pipeline to evaluate the consistency and reliability of clustering results by leveraging machine learning classification and prediction methods, with the goal of maintaining spatial proximity and gene expression patterns within clusters. In addition, it enables iterative correction of misclassified cells, providing actionable guidance for cluster refinement. We benchmarked STEAM on various public datasets, spanning multicell to single-cell resolution, normal and diseased tissues, as well as spatial transcriptomics and proteomics. The results highlighted its robustness and generalizability through comprehensive statistical evaluation metrics, such as Kappa score, F1 score, accuracy, adjusted rand index, normalized mutual information, and percentage of abnormal spots. Notably, STEAM supports multisample training, enabling cross-replicate clustering consistency assessments. Moreover, STEAM provides practical guidance by comparing clustering results across multiple approaches; here, we evaluated four representative methods encompassing both spatial-aware and spatial-ignorant strategies. In summary, STEAM is a promising tool for evaluating clustering robustness and benchmarking clustering performance for spatial omics data, offering valuable insights to drive reproducible discoveries in spatial biology.
Keywords:
spatial transcriptomics
cluster benchmarking
computational omics pipeline
classification and prediction
Journal
IF:
7.7
Papers:
5.6K
Citations:
2.7W


