返回
An Algorithm for Clustering Categorical Data With Set-Valued Features
DOI:10.1109/TNNLS.2017.2770167.png)
摘要
En 中文
In data mining, objects are often represented by a set of features, where each feature of an object has only one value. However, in reality, some features can take on multiple values, for instance, a person with several job titles, hobbies, and email addresses. These features can be referred to as set-valued features and are often treated with dummy features when using existing data mining algorithms to analyze data with set-valued features. In this paper, we propose an SV-k-modes algorithm that clusters categorical data with set-valued features. In this algorithm, a distance function is defined between two objects with set-valued features, and a set-valued mode representation of cluster centers is proposed. We develop a heuristic method to update cluster centers in the iterative clustering process and an initialization algorithm to select the initial cluster centers. The convergence and complexity of the SV-k-modes algorithm are analyzed. Experiments are conducted on both synthetic data and real data from five different applications. The experimental results have shown that the SV-k-modes algorithm performs better when clustering real data than do three other categorical clustering algorithms and that the algorithm is scalable to large data.
Keyword:
Categorical data set-valued feature
set-valued modes
SV-k-modes algorithm
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
8.9
论文数:
7.6K
被引数:
7.2W
机构
引用论文
Stratified feature sampling method for ensemble clustering of high dimensional data高维数据集成聚类的分层特征抽样方法
PATTERN RECOGNITION
IF7.6

