arrow
返回

Constructing bi-plots for random forest: Tutorial

delete2020-09-01
delete49
delete
OA
AI
L
Lionel Blanchet
R
Raffaele Vitale
R
Robert van Vorstenbosch
G
George Stavropoulos
J
John Pender
D
Daisy Jonkers
F
Frederik‐Jan van Schooten
A
Agnieszka Smolinska *
DOI:10.1016/j.aca.2020.06.043delete
delete原文链接
delete分享
delete收藏
查看原文
摘要

摘要

En 中文
Current technological developments have allowed for a significant increase and availability of data. Consequently, this has opened enormous opportunities for the machine learning and data science field, translating into the development of new algorithms in a wide range of applications in medical, biomedical, daily-life, and national security areas. Ensemble techniques are among the pillars of the machine learning field, and they can be defined as approaches in which multiple, complex, independent/uncorrelated, predictive models are subsequently combined by either averaging or voting to yield a higher model performance. Random forest (RF), a popular ensemble method, has been successfully applied in various domains due to its ability to build predictive models with high certainty and little necessity of model optimization. RF provides both a predictive model and an estimation of the variable importance. However, the estimation of the variable importance is based on thousands of trees, and therefore, it does not specify which variable is important for which sample group. The present study demonstrates an approach based on the pseudo-sample principle that allows for construction of bi-plots (i.e. spin plots) associated with RF models. The pseudo-sample principle for RF. is explained and demonstrated by using two simulated datasets, and three different types of real data, which include political sciences, food chemistry and the human microbiome data. The pseudo-sample bi plots, associated with RF and its unsupervised version, allow for a versatile visualization of multivariate models, and the variable importance and the relation among them. (c) 2020 Elsevier B.V. All rights reserved.
Keyword:
Random forest interpretation
Pseudo samples
Bi-plots
Proximity matrix
Principal coordinates analysis
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

Analytica Chimica Acta 封面图
Analytica Chimica Acta
IF:
6
论文数:
3.3W
被引数:
6.1W

机构

M
Maastricht University
学者数:
3.1W
论文数: 2.8W
被引数: 277
C
cnrs - institute of chemistry (inc)
学者数:
1.7W
论文数: 1.3W
被引数: 19
U
universite de lille
学者数:
2.7W
论文数: 2.0W
被引数: 15
学者 查看更多机构
引用论文

引用论文

A novel HLA‐B44 allele, B*44:127, was identified by sequencing‐based typing
err2011-07-31
err0
PREAI
errS.‐X. Jiao; X.‐Y. Chi; B. Hu; Z.‐H. Feng; L. Zhao
err分享
err收藏
An overview of the compactness of a range of axial flux PM machines
err2015-05-01
err0
PREAI
errD. J. Patterson; G. Heins; M. Turner; B. J. Kennedy; M. D. Smith; R. Rohoza
err分享
err收藏
Volatile Organic Compounds of Whole‐Grain Soft Winter Wheat
err2017-04-12
err0
PREAI
errTaehyun Ji; Moonseok Kang; Byung‐Kee Baik
err分享
err收藏
err分享
err收藏
Representative subset selection代表性子集选择
err2002-09-01
err279
PREAI
errDaszykowski, M; Walczak, B; Massart, DL
err分享
err收藏
The fecal microbiota as a biomarker for disease activity in Crohn's disease
err2016-10-13
err58
errOAAI
errTedjo, Danyta. I.; Smolinska, Agnieszka; Savelkoul, Paul H.; Masclee, Ad A.; van Schooten, Frederik J.; Pierik, Marieke J.; Penders, John; Jonkers, Daisy M. A. E.
err分享
err收藏
The Composition of Odor Compounds Emitted from Municipal Solid Waste Landfill
err2007-12-31
err0
errOAAI
errYoun-Suk Son; Jo-Chun Kim; Ki-Hyung Kim; Bo-A Lim; Kang-Nam Park; Woo-Keun Lee
err分享
err收藏
学者 查看更多内容