返回
An instance level analysis of data complexity
DOI:10.1007/s10994-013-5422-z.png)
摘要
En 中文
Most data complexity studies have focused on characterizing the complexity of the entire data set and do not provide information about individual instances. Knowing which instances are misclassified and understanding why they are misclassified and how they contribute to data set complexity can improve the learning process and could guide the future development of learning algorithms and data analysis methods. The goal of this paper is to better understand the data used in machine learning problems by identifying and analyzing the instances that are frequently misclassified by learning algorithms that have shown utility to date and are commonly used in practice. We identify instances that are hard to classify correctly (instance hardness) by classifying over 190,000 instances from 64 data sets with 9 learning algorithms. We then use a set of hardness measures to understand why some instances are harder to classify correctly than others. We find that class overlap is a principal contributor to instance hardness. We seek to integrate this information into the training process to alleviate the effects of class overlap and present ways that instance hardness can be used to improve learning.
Keyword:
Instance hardness
Dataset hardness
Data complexity
期刊
IF:
2.9
论文数:
2.7K
被引数:
3.4W
机构
引用论文
Physiological and behavioral responses to corticotropin-releasing factor administration: is CRF a mediator of anxiety or stress responses?促肾上腺皮质激素释放因子的生理和行为反应: CRF是焦虑或应激反应的中介物吗?

