Return
Clustering methods for high-dimensional data
DOI:10.5351/KJAS.2025.38.5.693.png)
Abstract
En 中文
High-dimensional data which is characterized by observations with thousands of features are prevalent in various fields such as gene expression data analysis, image processing, and natural language processing, etc. Despite of offering rich information, such data present substantial challenges for clustering due to the curse of dimensionality, degradation of similarity measures, and lack of interpretability. In response, a wide array of methodologies has been developed so far, including subspace clustering, dimension reduction, model-based approaches, and feature selection, etc. More recent advances incorporate regularization techniques and bootstrap-based procedures for estimating the number of clusters. This paper provides a comprehensive overview of contemporary clustering methods tailored for high-dimensional data, critically evaluating their strengths, weaknesses, and applicability. Furthermore, it outlines promising research directions aimed at developing efficient, robust, scalable and interpretable clustering algorithms.
Keywords:
cluster analysis
dimension reduction
feature selection
high-dimensional data
model-based clustering
subspace clustering
Journal
K
IF:
0
Papers:
20
Citations:
0

