Return
Large Scale Document Categorization With Fuzzy Clustering
DOI:10.1109/TFUZZ.2016.2604009.png)
Abstract
En 中文
Clustering documents into coherent categories is a very useful and important step for document processing and understanding. The introducing of fuzzy set theory into clustering provides a favorable mechanism to capture overlapping among document clusters. Document dataset is commonly represented as a collection of high-dimensional vectors, which may not be able to fit into memory entirely, when the dataset is large and with a very high dimensionality. However, most of the existing fuzzy clustering approaches deal with small and static datasets. Some of them may have a good scalability but they are only effective for low dimensional data. The study presented in this paper is about new efforts on fuzzy clustering of large-scale and high-dimensional data-especially suitable for document categorization. To consider both large scale and high dimensionality into the problem formulation, our key idea is to incorporate document-tailored fuzzy clustering into a scheme, which is effective for dealing with a large-scale problem. We first identified three representative schemes in fuzzy clustering for handling large-scale data, namely sampling extension, single pass, and divide ensemble. The limitation of fuzzy C-means (FCM)-based approaches for a large document clustering are then investigated. Based on the study, we propose new approaches by incorporating each of hyperspherical FCM and fuzzy coclustering with the three scale-up schemes, respectively. This enables our new approaches to maintain effectiveness for high-dimensional data with an extended scalability. Extensive experimental studies with real-world large document datasets have been conducted and the results demonstrate that the proposed approaches perform consistently better over existing ones in document categorization.
Keywords:
Document clustering
fuzzy clustering
fuzzy coclustering (FCoC)
hyperspherical clustering
large data
AI Summary
Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.
Journal
IF:
11.9
Papers:
4.9K
Citations:
2.9W

