返回
Pattern Learning and Knowledge Distillation for Single-Cell Data Annotation
DOI:10.3390/biology15010002.png)
摘要
En 中文
Transferring cell type annotations from reference dataset to query dataset is a fundamental problem in AI-based single-cell data analysis. However, single-cell measurement techniques lead to domain gaps between multiple batches or datasets. The existing deep learning methods lack consideration on batch integration when learning reference annotations, which is a challenge for cell type annotation on multiple query batches. For cell representation, batch integration can not only eliminate the gaps between batches or datasets but also improve the heterogeneity of cell clusters. In this study, we proposed PLKD, a cell type annotation method based on pattern learning and knowledge distillation. PLKD consists of Teacher (Transformer) and Student (MLP). Teacher groups all input genes (features) into different gene sets (patterns), and each pattern represents a specific biological function. This design enables model to focus on biologically relevant functions interaction rather than gene-level expression that is susceptible to gaps of batches. In addition, knowledge distillation makes lightweight Student resistant to noise, allowing Student to infer quickly and robustly. Furthermore, PLKD supports multi-modal cell type annotation, multi-modal integration and other tasks. Benchmark experiments demonstrate that PLKD is able to achieve accurate and robust cell type annotation.
Keyword:
pattern learning
knowledge distillation
cell type annotation
batch integration
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
B
IF:
3.5
论文数:
709
被引数:
0
机构
引用论文
Transcriptome-scale spatial gene expression in the human dorsolateral prefrontal cortex人背外侧前额叶皮层中转录组尺度的空间基因表达
High-throughput sequencing of the transcriptome and chromatin accessibility in the same cell
NATURE BIOTECHNOLOGY
IF41.7
Modal-NexT: Towards unified heterogeneous cellular data integrationModal-NexT: 走向统一的异构蜂窝数据集成
Information Fusion
IF15.5
The scverse project provides a computational ecosystem for single-cell omics data analysisscverse项目为单细胞组学数据分析提供了一个计算生态系统
NATURE BIOTECHNOLOGY
IF41.7
Modal-nexus auto-encoder for multi-modality cellular data integration and imputation
NATURE COMMUNICATIONS
IF15.7

