返回
A fuzzy k-prototype clustering algorithm for mixed numeric and categorical data
DOI:10.1016/j.knosys.2012.01.006.png)
摘要
En 中文
In many applications, data objects are described by both numeric and categorical features. The k-prototype algorithm is one of the most important algorithms for clustering this type of data. However, this method performs hard partition, which may lead to misclassification for the data objects in the boundaries of regions, and the dissimilarity measure only uses the user-given parameter for adjusting the significance of attribute. In this paper, first, we combine mean and fuzzy centroid to represent the prototype of a cluster, and employ a new measure based on co-occurrence of values to evaluate the dissimilarity between data objects and prototypes of clusters. This measure also takes into account the significance of different attributes towards the clustering process. Then we present our algorithm for clustering mixed data. Finally, the performance of the proposed method is demonstrated by a series of experiments on four real world datasets in comparison with that of traditional clustering algorithms. (C) 2012 Elsevier B.V. All rights reserved.
Keyword:
Fuzzy clustering
Data mining
Mixed data
Dissimilarity measure
Attribute significance
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
K
IF:
7.6
论文数:
1.3W
被引数:
4.5W
机构
引用论文
The Inhibitory Effect of Cordycepin on the Proliferation of MCF-7 Breast Cancer Cells, and Its Mechanism: An Investigation Using Network Pharmacology-Based Analysis
Biomolecules
IF0
G-ANMI: A mutual information based genetic clustering algorithm for categorical dataG-anmi: 一种基于互信息的分类数据遗传聚类算法

