返回
Self-Tuned Descriptive Document Clustering Using a Predictive Network
DOI:10.1109/TKDE.2017.2781721.png)
摘要
En 中文
Descriptive clustering consists of automatically organizing data instances into clusters and generating a descriptive summary for each cluster. The description should inform a user about the contents of each cluster without further examination of the specific instances, enabling a user to rapidly scan for relevant clusters. Selection of descriptions often relies on heuristic criteria. We model descriptive clustering as an auto-encoder network that predicts features from cluster assignments and predicts cluster assignments from a subset of features. The subset of features used for predicting a cluster serves as its description. For text documents, the occurrence or count of words, phrases, or other attributes provides a sparse feature representation with interpretable feature labels. In the proposed network, cluster predictions are made using logistic regression models, and feature predictions rely on logistic or multinomial regression models. Optimizing these models leads to a completely self-tuned descriptive clustering approach that automatically selects the number of clusters and the number of features for each cluster. We applied the methodology to a variety of short text documents and showed that the selected clustering, as evidenced by the selected feature subsets, are associated with a meaningful topical organization.
Keyword:
Descriptive clustering
feature selection
logistic regression
model selection
sparse models
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
10.4
论文数:
6.8K
被引数:
3.2W
机构
引用论文
Coordination complexes of niobium and tantalum. V. Eight-coordinated di- and triperoxoniobates(V) and -tantalates(V) with some nitrogen and oxygen bidentate ligands铌和钽的配合物。五。具有一些氮和氧双齿配体的八配位的二过氧化物和三过氧化物铌酸盐 (V) 和钽酸盐 (V)
IF0

