arrow
返回

Predicting sample size required for classification performance

delete2012-02-15
delete420
delete
OA
AI
R
Rosa L. Figueroa
Q
Qing Zeng‐Treitler *
S
Sasikiran Kandula
L
Long Ngo
DOI:10.1186/1472-6947-12-8delete
delete原文链接
delete分享
delete收藏
查看原文
摘要

摘要

En 中文
Background: Supervised learning methods need annotated data in order to generate efficient models. Annotated data, however, is a relatively scarce resource and can be expensive to obtain. For both passive and active learning methods, there is a need to estimate the size of the annotated sample required to reach a performance target. Methods: We designed and implemented a method that fits an inverse power law model to points of a given learning curve created using a small annotated training set. Fitting is carried out using nonlinear weighted least squares optimization. The fitted model is then used to predict the classifier's performance and confidence interval for larger sample sizes. For evaluation, the nonlinear weighted curve fitting method was applied to a set of learning curves generated using clinical text and waveform classification tasks with active and passive sampling methods, and predictions were validated using standard goodness of fit measures. As control we used an un-weighted fitting method. Results: A total of 568 models were fitted and the model predictions were compared with the observed performances. Depending on the data set and sampling method, it took between 80 to 560 annotated samples to achieve mean average and root mean squared error below 0.01. Results also show that our weighted fitting method outperformed the baseline un-weighted method (p < 0.05). Conclusions: This paper describes a simple and effective sample size prediction algorithm that conducts weighted fitting of learning curves. The algorithm outperformed an un-weighted algorithm described in previous literature. It can help researchers determine annotation sample size for supervised machine learning.
Keyword:
POWER
ACCURACY
NUMBER
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

BMC Medical Informatics and Decision Making 封面图
BMC Medical Informatics and Decision Making
IF:
3.8
论文数:
4.4K
被引数:
1.2W

机构

U
University of Utah
学者数:
3.0W
论文数: 2.2W
被引数: 4.6W
U
Utah System of Higher Education
学者数:
4.6W
论文数: 4.0W
被引数: 161
U
universidad de concepcion
学者数:
8.4K
论文数: 6.4K
被引数: 8
学者 查看更多机构
引用论文

引用论文

How large a training set is needed to develop a classifier for microarray data?
err2008-01-02
err104
errOAAI
errDobbin, Kevin K.; Zhao, Yingdong; Simon, Richard M.
err分享
err收藏
On the nature of Upsilon Sagittarii
err1983-05-01
err0
errOAAI
errD. Schoenberner; J. S. Drilling
err分享
err收藏
err
IF0
err
err0
PREAI
err
err分享
err收藏
Ifosfamide in pediatric malignant solid tumors
err1989-07-01
err0
PREAI
errC.B. Pratt; E.C. Douglass; E.L. Etcubanas; M.P. Goren; A.A. Green; F.A. Hayes; M.E. Horowitz; W.H. Meyer; E.I. Thompson; J.A. Wilimas
err分享
err收藏
The RNA binding proteins RBM38 and DND1 are repressed in AML and have a novel function in APL differentiation
err2016-02-01
err0
PREAI
errJulian Wampfler; Elena A. Federzoni; Bruce E. Torbett; Martin F. Fey; Mario P. Tschan
err分享
err收藏
学者 查看更多内容