返回
A principal feature analysis
DOI:10.1016/j.jocs.2021.101502.png)
摘要
En 中文
A key task of data science is to identify relevant features linked to certain output variables that are supposed to be modeled or predicted. To obtain a small but meaningful model, it is important to find stochastically independent variables capturing all the information necessary to model or predict the output variables sufficiently. Therefore, we introduce in this work a framework to detect linear and non-linear dependencies between different features. As we will show, features that are actually functions of other features do not represent further information. Consequently, a model reduction neglecting such features conserves the relevant information, reduces noise and thus improves the quality of the model. Furthermore, a smaller model makes it easier to adopt a model of a given system. In addition, the approach structures dependencies within all the considered features. This provides advantages for classical modeling starting from regression ranging to differential equations and for machine learning. To show the generality and applicability of the presented framework 2154 features of a data center are measured and a model for classifying faulty and non-faulty states of the data center is set up. This number of features is automatically reduced by the framework to 161 features. The prediction accuracy for the reduced model even improves compared to the model trained on the total number of features. A second example is the analysis of a gene expression data set where from 9513 genes 9 genes are extracted from whose expression levels two cell clusters of macrophages can be distinguished.
Keyword:
Feature selection
Classify variables into functions and arguments
Combine statistics and graph theory
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
18.3
论文数:
3.1K
被引数:
4.0K
机构
引用论文
Integrating single-cell transcriptomic data across different conditions, technologies, and species跨不同条件、技术和物种整合单细胞转录组数据
NATURE BIOTECHNOLOGY
IF41.7
Facile synthesis of tremelliform Co0.85Se nanosheets: An efficient catalyst for the decomposition of hydrazine hydratetremelliform Co0.85Se纳米片的简易合成: 水合肼分解的有效催化剂

