arrow
返回

A preprocess algorithm of filtering irrelevant information based on the minimum class difference

delete2006-10-01
delete9
PRE
AI
Z
Zhiping Chen
K
Kevin Lü *
DOI:10.1016/j.knosys.2006.03.005delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Whether a word (or a feature) should be included or excluded during the process of text classification could depend on a number of factors, such as the amount of information it represents, its appearance frequency and its meaning. The application context is another important factor that needs to be considered. A word may be able to represent the characteristic of a document in one application context but may not reflect its nature in another. This paper reports on an investigation into the selection of features for classification with the consideration of the application context of the documents to be processed. A new feature selection algorithm for text classification to be known as the PBMCD algorithm is proposed. This algorithm has been implemented and tested using three different data sets. The experiment results have shown that this algorithm cannot only filter out irrelevant features before the classification process but also can increase the classification accuracy. As a comparison, experiment results with other methods have also been presented. (c) 2006 Elsevier B.V. All rights reserved.
Keyword:
classification
text categorization
feature selection
preprocess
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

K
Knowledge-Based Systems
IF:
7.6
论文数:
1.2W
被引数:
4.5W

机构

暂无机构信息
引用论文

引用论文

Use of Adsol® preservation solution for prolonged storage of low viscosity AS‐1 red blood cells
err2008-07-07
err0
PREAI
errA. Heaton; J. Miripol; R. Aster; P. Hartman; D. Dehart; L. Rzad; B. Grapka; W. Davisson; D. H. Buchholz
err分享
err收藏
err分享
err收藏