arrow
返回

Simple instance selection for bankruptcy prediction

delete2012-03-01
delete50
PRE
AI
C
Chih‐Fong Tsai *
K
Kai-Chun Cheng
DOI:10.1016/j.knosys.2011.09.017delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Instance selection or outlier detection is an important task during data mining, which focuses on filtering out bad data from a given dataset. However, there is no rigid mathematical definition of what constitutes an outlier and an outlier is not a binary property. Therefore, different volumes of outliers may be detected depending on the setting of the threshold for what constitutes an outlier, e.g., the distance in distance-based outlier detection. In this study, we examine bankruptcy prediction performance achieved after removal of different outlier volumes from four widely used datasets, namely the Australian, German, Japanese, and UC Competition datasets. Specifically, a simple distance-based clustering outlier detection method is used. In addition, four popular classification techniques are compared, artificial neural networks, decision trees, logistic regression, and support vector machines. Experiments are conducted to examine (1) the prediction performance of the bankruptcy prediction models with and without instance selection, (2) the stability of bankruptcy prediction models after the removal of outliers from the testing set, and (3) the characteristics of these four different datasets. The results show that with the German dataset it is much more difficult for the prediction models to provide high rates of accuracy after outlier removal, while it is easier with the UC Competition dataset. Removing 50% of the outliers can lead to optimal performance of these four models. In addition, using the removed outliers to test the prediction accuracy of these models, we find that it is support vector machines (SVM) that provide the highest rate of prediction accuracy and perform with much more stability and good noise tolerance than the other three prediction models. Furthermore, the prediction accuracy of the SVM model followed by instance selection is similar to the one without instance selection (i.e., the SVM baseline). In other words, the difference in performance between the SVM and the SVM baseline is the least of the three models in comparison with their corresponding baselines. (C) 2011 Elsevier B.V. All rights reserved.
Keyword:
Instance selection
Outlier detection
Data mining
Bankruptcy prediction
Clustering
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

K
Knowledge-Based Systems
IF:
7.6
论文数:
1.3W
被引数:
4.5W

机构

N
National Central University
学者数:
1.0W
论文数: 8.6K
被引数: 6.4K
引用论文

引用论文

err分享
err收藏
err分享
err收藏
Fast wrapper feature subset selection in high-dimensional datasets by means of filter re-ranking
err2012-02-01
err123
PREAI
errBermejo, Pablo; de la Ossa, Luis; Gamez, Jose A.; Puerta, Jose M.
err分享
err收藏
IoT Based Design of Air Quality Monitoring System Web Server for Android Platform
err2021-02-09
err0
PREAI
errKoel Datta Purkayastha; Ritesh Kishore Mishra; Arunava Shil; Sambhu Nath Pradhan
err分享
err收藏
The random subspace binary logit (RSBL) model for bankruptcy prediction
err2011-12-01
err48
PREAI
errLi, Hui; Lee, Young-Chan; Zhou, Yan-Chun; Sun, Jie
err分享
err收藏
学者 查看更多内容