arrow
Return

A simple and effective outlier detection algorithm for categorical data

delete2013-09-27
delete22
PRE
AI
X
Xingwang Zhao *
J
Jiye Liang
F
Fuyuan Cao
DOI:10.1007/s13042-013-0202-4delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Outlier detection is an important data mining task that has attracted substantial attention within diverse research communities and the areas of application. By now, many techniques have been developed to detect outliers. However, most existing research focus on numerical data. And they can not directly apply to categorical data because of the difficulty of defining a meaningful similarity measure for categorical data. In this paper, a weighted density definition is given firstly, which takes account of the density and uncertainty of objects in every attributes simultaneously. Furthermore, a simple and effective outlier detection algorithm for categorical data based on the given weighted density is proposed. The corresponding time complexity of the algorithm is analyzed as well. Experimental results on real and synthetic data sets demonstrate the effectiveness and efficiency of our proposed algorithm.
Keywords:
Outlier detection
Categorical data
Weighted density
Information entropy

Journal

International Journal of Machine Learning and Cybernetics cover
International Journal of Machine Learning and Cybernetics
IF:
2.7
Papers:
3.1K
Citations:
5.6K

Organization

S
Shanxi University
Scholars:
1.3W
Papers: 8.4K
Citations: 1.2W