arrow
Return

Characterizing efficient feature selection for single-cell expression analysis

delete2024-07-08
delete1
delete
OA
AI
J
Juok Cho
B
Bukyung Baik
N
Nguyen Cao Truong Hai
D
Daeui Park
D
Dougu Nam *
DOI:10.1093/bib/bbae317delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
Unsupervised feature selection is a critical step for efficient and accurate analysis of single-cell RNA-seq data. Previous benchmarks used two different criteria to compare feature selection methods: (i) proportion of ground-truth marker genes included in the selected features and (ii) accuracy of cell clustering using ground-truth cell types. Here, we systematically compare the performance of 11 feature selection methods for both criteria. We first demonstrate the discordance between these criteria and suggest using the latter. We then compare the distribution of selected genes in their means between feature selection methods. We show that lowly expressed genes exhibit seriously high coefficients of variation and are mostly excluded by high-performance methods. In particular, high-deviation- and high-expression-based methods outperform the widely used in Seurat package in clustering cells and data visualization. We further show they also enable a clear separation of the same cell type from different tissues as well as accurate estimation of cell trajectories.
Keywords:
single-cell RNA-sequencing
feature selection
clustering
trajectory analysis
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

Briefings in Bioinformatics cover
Briefings in Bioinformatics
IF:
7.7
Papers:
5.6K
Citations:
2.7W

Organization

K
Korea Institute of Toxicology
Scholars:
995
Papers: 760
Citations: 934