返回
Patch clustering for massive data sets
DOI:10.1016/j.neucom.2008.12.026.png)
摘要
En 中文
The presence of huge data sets poses new problems to popular clustering and visualization algorithms such as neural gas (NG) and the self-organising-map (SOM) due to memory and time constraints. In such situations, it is no longer possible to store all data points in the main memory at once and only a few, ideally only one run over the whole data set is still affordable to achieve a feasible training time. In this contribution we propose single pass extensions of the classical clustering algorithms NG and SOM which are based on a simple patch decomposition of the data set and fast batch optimization schemes of the underlying cost function. The algorithms only require a fixed memory space. They maintain the benefits of the original ones including easy implementation and interpretation as well as large flexibility and adaptability. We demonstrate that parallelization of the methods becomes easily possible and we show the efficiency of the approach in a variety of experiments. (C) 2009 Elsevier B.V. All rights reserved.
Keyword:
Neural gas
Clustering streaming data
Parallelization
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
6.5
论文数:
2.5W
被引数:
6.5W
机构
引用论文
Motif-Directed Oxidative Folding to Design and Discover Multicyclic Peptides for Protein RecognitionMotif-导向氧化折叠法设计与发现用于蛋白质识别的多环肽
Decision templates for multiple classifier fusion: an experimental comparison
PATTERN RECOGNITION
IF7.6

