Return
A novel clustering algorithm for categorical data with MGR based reference set selection method
DOI:10.1016/j.neucom.2025.132003.png)
Abstract
En 中文
It is difficult for clustering to measure the similarity or dissimilarity between two categorical data objects because the categorical data lack of a clear space structure. Some researchers proposed methods to map categorical objects into Euclidean space to enhance the distinguishability of objects. However, the existing spatial mapping methods and related categorical data clustering algorithms only consider the object level and do not exploit the attributes of the categorical data to reduce the datasets. And when the dataset is large, the current algorithms are very time-consuming. Categorical data attributes inherently partition dataset. Therefore, it is necessary to analyse data distributions at the attribute level in order to select reference sets that more appropriately represent the distribution of categorical data and to construct the space structure of categorical data. In this paper, a new clustering algorithm for categorical data with Mean Gain Ratio (MGR) based reference set selection method is proposed. In detail, firstly, a MGR based method is given for selecting a more appropriate reference set. This method first selects the attribute with the highest MGR, and then selects an object from each equivalent class of the partition generated by that attribute to form a reference set. And then, we present a clustering algorithm for categorical data by combing the proposed MGR based representation and the -means algorithm. The results of comparative experiments show that the proposed method enhances higher clustering performance than existing methods. Furthermore, when facing up to the very large datasets, the algorithm proposed in this paper is better in terms of time complexity and scalability.
Journal
IF:
6.5
Papers:
2.5W
Citations:
6.5W
Organization
No organization information available

