返回
A randomized algorithm for clustering discrete sequences
DOI:10.1016/j.patcog.2024.110388.png)
摘要
En 中文
Cluster analysis is one of the most important research issues in data mining and machine learning. To date, numerous clustering algorithms have been proposed to tackle the fixed -length vector data. In many real applications, we need to detect clusters from a set of discrete sequences in which each sequence is an ordered list of items. Due to the sequential and discrete nature, the discrete sequence clustering problem is more challenging and most of existing vector data clustering algorithms cannot be directly employed. In this paper, we present a stochastic algorithm for clustering discrete sequences. Our method first quickly generates a set of random partitions over the sequential data set and then merges these random clustering results via weighted graph construction and partition. We perform extensive empirical comparisons on real data sets to show that our method is comparable to those state-of-the-art clustering algorithms with respect to both accuracy and efficiency.
Keyword:
Sequence clustering
Sequential data analysis
Cluster analysis
Randomized algorithm
期刊
IF:
7.6
论文数:
1.3W
被引数:
4.5W
机构
引用论文
Transmission and Scanning Electronmicroscopy Study of the Action of Sage and Rosemary Essential Oils and Eucalyptol on Candida albicans/Transmissions‐ und rasterelektronenmikroskopische Untersuchungen zur Wirkung von Salbeiöl, Rosmarinöl und Eucalyptol auf Candida albicans
Mycoses
IF0

