arrow
Return

FORESTEXTER: An efficient random forest algorithm for imbalanced text categorization

delete2014-09-01
delete100
PRE
AI
吴
吴庆耀 (Qingyao Wu) *
Y
Yunming Ye
H
Haijun Zhang
M
Michael K. Ng
S
Shen-Shyang Ho
DOI:10.1016/j.knosys.2014.06.004delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
In this paper, we propose a new random forest (RF) based ensemble method, FORESTMER, to solve the imbalanced text categorization problems. RF has shown great success in many real-world applications. However, the problem of learning from text data with class imbalance is a relatively new challenge that needs to be addressed. A RF algorithm tends to use a simple random sampling of features in building their decision trees. As a result, it selects many subspaces that contain few, if any, informative features for the minority class. Furthermore, the Gini measure for data splitting is considered to be skew sensitive and bias towards the majority class. Due to the inherent complex characteristics of imbalanced text datasets, learning RF from such data requires new approaches to overcome challenges related to feature subspace selection and cut-point choice while performing node splitting. To this end, we propose a new tree induction method that selects splits, both feature subspace selection and splitting criterion, for RF on imbalanced text data. The key idea is to stratify features into two groups and to generate effective term weighting for the features. One group contains positive features for the minority class and the other one contains the negative features for the majority class. Then, for feature subspace selection, we effectively select features from each group based on the term weights. The advantage of our approach is that each subspace contains adequate informative features for both minority and majority classes. One difference between our proposed tree induction method and the classical RF method is that our method uses Support Vector Machines (SVM) classifier to split the training data into smaller and more balance subsets at each tree node, and then successively retrains the SVM classifiers on the data partitions to refine the model while moving down the tree. In this way, we force the classifiers to learn from refined feature subspaces and data subsets to fit the imbalanced data better. Hence, the tree model becomes more robust for text categorization task with imbalanced dataset. Experimental results on various benchmark imbalanced text datasets (Reuters-21578, Ohsumed, and imbalanced 20 newsgroup) consistently demonstrate the effectiveness of our proposed FORESTEXTER method. The performance of our proposed approach is competitive against the standard random forest and different variants of SVM algorithms. (C) 2014 Elsevier B.V. All rights reserved.
Keywords:
Text categorization
Imbalanced classification
Random forests
SVM
Stratified sampling
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

K
Knowledge-Based Systems
IF:
7.6
Papers:
1.3W
Citations:
4.5W

Organization

H
harbin institute of technology
Scholars:
8.0W
Papers: 6.6W
Citations: 66
H
Hong Kong Baptist University
Scholars:
6.3K
Papers: 7.5K
Citations: 1.3W
N
Nanyang Technological University
Scholars:
4.9W
Papers: 4.8W
Citations: 8.1W
researcher View more organizations
Cited Papers

Cited Papers

Semantic search in the World News domain using automatically extracted metadata files
err2012-03-01
err15
PREAI
errKallipolitis, Leonidas; Karpis, Vassilis; Karali, Isambo
errShare
errSave
Classical dynamical theory of heavy ion fusion and scattering
err1974-12-01
err0
PREAI
errJ.P. Bondorf; M.I. Sobel; D. Sperber
errShare
errSave
Solid awakening
err2008-02-20
err0
errOAAI
errLeonard R. MacGillivray
errShare
errSave
Effects of anticholinesterase drugs on biomarkers and behavior of pumpkinseed, Lepomis gibbosus (Linnaeus, 1758)
err2012-01-01
err0
PREAI
errSara Rodrigues; Sara C. Antunes; Fátima P. Brandão; Bruno B. Castro; Fernando Gonçalves; Bruno Nunes
errShare
errSave
Arsenic removal from aqueous solutions by adsorption using novel MIL-53(Fe) as a highly efficient adsorbent
err2015-01-01
err0
PREAI
errTuan. A. Vu; Giang. H. Le; Canh. D. Dao; Lan. Q. Dang; Kien. T. Nguyen; Quang. K. Nguyen; Phuong. T. Dang; Hoa. T. K. Tran; Quang. T. Duong; Tuyen. V. Nguyen; Gun. D. Lee
errShare
errSave
errShare
errSave
researcher View more