返回
Finding significant keywords for document databases by two-phase Maximum Entropy Partitioning
DOI:10.1016/j.patrec.2019.04.023.png)
摘要
En 中文
This paper investigates the selection of class-specific significant keywords for document databases. We define two types of significant keywords with respect to a document class: Elite and Unique Elite, derived in two phases. Elite Keywords are defined as those that have high term frequencies within the class. To obtain the top partition of distinctively high occurring terms in each class, we employ Maximum Entropy Partitioning (MEP) in the first phase. Our presumption is that the term probabilities within the subset of significant (and non-significant) keywords at the point of maximum entropy are relatively more uniform with respect to each other. Unique Elite keywords are those that are Elite for a particular class, and at the same time have a higher frequency of occurrence only in that class as compared to the other classes. To measure this aspect, in the second phase, we compute the entropy of each Elite keyword across all classes, sort the entropies in the ascending order and again employ MEP to shortlist those Elite keywords that occur uniquely in this class, characterized by distinctively low entropy. Experimental comparisons with the state-of-the-art on benchmark datasets using an ensemble of bagged tree classifiers, establishes the discriminatory powers of the derived keywords. (C) 2019 Elsevier B.V. All rights reserved.
Keyword:
Document categorization
Keyword extraction
Maximum Entropy Partitioning
Unique Elite keywords
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
3.3
论文数:
7.9K
被引数:
1.6W
机构
引用论文
An experimental comparison of three methods for constructing ensembles of decision trees: Bagging, boosting, and randomization
MACHINE LEARNING
IF2.9

