arrow
返回

Comparing software prediction techniques using simulation

delete2001-01-01
delete164
delete
OA
AI
M
Martin Shepperd *
G
Gada Kadoda
DOI:10.1109/32.965341delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
The need for accurate software prediction systems increases as software becomes much larger and more complex. A variety of techniques have been proposed; however, none has proven consistently accurate and there is still much uncertainty as to what technique suits which type of prediction problem. We believe that the underlying characteristics-size, number of features, type of distribution, etc-of the data set influence the choice of the prediction system to be used. For this reason, we would like to control the characteristics of such data sets in order to systematically explore the relationship between accuracy, choice of prediction system, and data set characteristic. Also, in previous work, it has proven difficult to obtain significant results over small data sets. Consequently, it would be useful to have a large validation data set. Our solution is to simulate data allowing both control and the possibility of large (1,000) validation cases. In this paper, we compared four prediction techniques: regression, rule induction, nearest neighbor (a form of case-based reasoning), and neural nets. The results suggest that there are significant differences depending upon the characteristics of the data set. Consequently, researchers should consider prediction context when evaluating competing prediction systems. We also observed that the more messy the data and the more complex the relationship with the dependent variable, the more variability in the results. In the more complex cases, we observed significantly different results depending upon the particular training set that has been sampled from the underlying data set. This suggests that researchers will need to exercise caution when comparing different approaches and utilize procedures such as bootstrapping in order to generate multiple samples for training purposes. However, our most important result is that it is more fruitful to ask which is the best prediction system in a particular context rather than which is the best prediction system.
Keyword:
prediction system
simulation
machine learning
data set characteristic

期刊

IEEE Transactions on Software Engineering 封面图
IEEE Transactions on Software Engineering
IF:
5.6
论文数:
2.8K
被引数:
1.1W

机构

暂无机构信息
引用论文

引用论文

Gender differences in social support and leisure-time physical activity
err2014-08-01
err0
errOAAI
errAldair J Oliveira; Claudia S Lopes; Mikael Rostila; Guilherme Loureiro Werneck; Rosane Härter Griep; Antônio Carlos Monteiro Ponce de Leon; Eduardo Faerstein
err分享
err收藏
err分享
err收藏
err
IF0
err
err0
PREAI
err
err分享
err收藏
学者 查看更多内容