返回
Classification Confidence in Exploratory Learning: A User's Guide
DOI:10.3390/make5030043.png)
摘要
En 中文
This paper investigates the post-hoc calibration of confidence for exploratory machine learning classification problems. The difficulty in these problems stems from the continuing desire to push the boundaries of which categories have enough examples to generalize from when curating datasets, and confusion regarding the validity of those categories. We argue that for such problems the one-versus-all approach (top-label calibration) must be used rather than the calibrate-the-full-response-matrix approach advocated elsewhere in the literature. We introduce and test four new algorithms designed to handle the idiosyncrasies of category-specific confidence estimation using only the test set and the final model. Chief among these methods is the use of kernel density ratios for confidence calibration including a novel algorithm for choosing the bandwidth. We test our claims and explore the limits of calibration on a bioinformatics application (PhANNs) as well as the classic MNIST benchmark. Finally, our analysis argues that post-hoc calibration should always be performed, may be performed using only the test dataset, and should be sanity-checked visually.
Keyword:
confidence calibration
top-label confidence calibration
bioinformatics
machine learning
exploratory machine learning
期刊
M
IF:
6
论文数:
834
被引数:
1.8K
机构
引用论文
Interactive CD34-Positive Fibroblasts and Factor XIIIa-Positive Histiocytes in Cutaneous Mesenchymal Tumors交互式CD34阳性成纤维细胞和Factor XIIIa阳性组织细胞在皮肤间叶组织肿瘤中
Studies on Lupin Alkaloids. IV. Total Syntheses of Optically Active Matrine and Allomatrine羽扇豆生物碱的研究。四。光学活性苦参碱和Allomatrine的总合成
没有更多内容

