返回
Clustering mixed-type data using a probabilistic distance algorithm
DOI:10.1016/j.asoc.2022.109704.png)
摘要
En 中文
Cluster analysis is a broadly used unsupervised data analysis technique for finding groups of homoge-neous units in a data set. Probabilistic distance clustering adjusted for cluster size (PDQ), discussed in this contribution, falls within the broad category of clustering methods initially developed to deal with continuous data; it has the advantage of fuzzy membership and robustness. However, a common issue in clustering deals with treating mixed-type data: continuous and categorical, which are among the most common types of data. This paper extends PDQ for mixed-type data using different dissimilarities for different kinds of variables. At first, the PDQ for mixed-type data is defined, then a simulation design shows its advantages compared to some state of the art techniques, and ultimately, it is used on a real data set. The conclusion includes some future developments.(c) 2022 The Authors. Published by Elsevier B.V. This is an open access article under the CC BY-NC-ND license (http://creativecommons.org/licenses/by-nc-nd/4.0/).
Keyword:
Probabilistic distance clustering
Mixed-type data
Fuzzy clustering
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
6.6
论文数:
1.4W
被引数:
4.8W
机构
引用论文
Heavy metal tolerance of marine phytoplankton. IV. Combined effect of zinc and cadmium on growth and uptake in some marine diatoms海洋浮游植物对重金属的耐受性。四。锌和镉对某些海洋硅藻生长和吸收的综合影响

