返回
A practical outlier detection approach for mixed-attribute data
DOI:10.1016/j.eswa.2015.07.018.png)
摘要
En 中文
Outlier detection in mixed-attribute space is a challenging problem for which only a few approaches have been proposed. However, such existing methods suffer from the fact that there is a lack of an automatic mechanism to formally discriminate between outliers and inliers. In fact, a common approach to outlier identification is to estimate an outlier score for each object and then provide a ranked list of points, expecting outliers to come first. A major problem of such an approach is where to stop reading the ranked list? How many points should be chosen as outliers? Other methods, instead of outlier ranking, implement various strategies that depend on user-specified thresholds to discriminate outliers from inliers. Ad-hoc threshold values are often used. With such an unprincipled approach it is impossible to be objective or consistent. To alleviate these problems, we propose a principled approach based on the bivariate beta mixture model to identify outliers in mixed-attribute data. The proposed approach is able to automatically discriminate outliers from inliers and it can be applied to both mixed-type attribute and single-type (numerical or categorical) attribute data without any feature transformation. Our experimental study demonstrates the suitability of the proposed approach in comparison to mainstream methods. (C) 2015 Elsevier Ltd. All rights reserved.
Keyword:
Data mining
Outlier detection
Mixed-attribute data
Mixture model
Bivariate beta
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
7.5
论文数:
3.0W
被引数:
10.2W
机构
暂无机构信息
引用论文
A fully Bayesian model based on reversible jump MCMC and finite Beta mixtures for clustering基于可逆跳MCMC和有限Beta混合的完全贝叶斯聚类模型
Outlier detection in relational data: A case study in geographical information systems关系数据中的异常值检测: 地理信息系统中的案例研究

